programmiere

Matrix

Ask a thousand pages a question. On a phone, offline.

Give it regulations, contracts or annual reports. Ask anything. You get an answer — and the exact sentence it came from, so you can check it in five seconds instead of trusting it.

One public annual report, turned into 4,385 entities and 1,919 relations. Click to see it full size.
Matrix knowledge graph extracted from one public annual report: thousands of entities connected by labelled relations.

The problem

AI sounds right. That's the problem.

Chatbots are fluent and confident — and can't tell you where it says that. For anyone whose job is to be right on paper, "sounds plausible" is worth nothing.

Regulations, contracts and reports are thousands of pages of scattered knowledge where every statement needs a source. Matrix is built for exactly that kind of reading.

The one rule

The AI reads. Everything you could check is plain code.

The language model is only used where language understanding is needed. Splitting documents, following references, verifying quotes, correcting sources and logging every step is done by ordinary rules — reproducible, no AI judgement involved.

Throughout this page: AI means the model decides,rules means deterministic code does.

Every answer shows its sources

Down to the verbatim quote in the original. Clickable, checkable.

Your data stays home

Runs with local models on the device or inside your own network. Nothing has to leave.

Nothing happens in the dark

Every search, every hit, every automatic correction is logged and can be replayed.

Ask a question

It doesn't guess. It goes looking.

Ask once, and the model researches the documents itself — like a junior analyst with a strict supervisor.

  1. Prepare the documents rules

    Split along their real structure — articles, sections. Each piece gets a "meaning fingerprint" so search works on sense, not just words. References like "pursuant to Art. 92(1)" are linked to their target.

  2. Search, read, search again AI

    The model writes its own search queries, opens passages, follows references and keeps digging with the vocabulary it found — until it has enough evidence or runs out of budget.

  3. Guard rails rules

    A search budget caps the cost. Watchdogs stop it from going in circles. An answer given without searching at all is thrown away.

  4. Answer only from what it found AI

    Every statement is tagged with its source number: [1], [2], [3].

  5. Check the answer rules

    The checks in the next section run on every answer, then everything is shown with clickable sources and a full log.

A real question to a public (German) annual report — “Which climate targets has Uniper committed to, and by when?” One sentence of answer, and right underneath the passages it came from, with their relevance scores. (Answered here by a hosted Llama model; local models run the same pipeline.)
Matrix chat: a question about a company's climate targets, a one-sentence answer, and the four source passages it was drawn from.

Built-in fact check

It checks its own homework.

After every answer, a chain of plain-code checks goes over each citation.

  1. Does the source exist?

    Every [n] is validated against the list of sources actually retrieved.

  2. Is the quote really there?

    Quoted words are searched in the original. Not found → a warning sign on the answer.

  3. Does that source actually say it?

    Two independent signals — shared key terms (legal references count triple) and closeness in meaning.

  4. Fix only when both agree

    Then the marker is pointed to the right source and labelled "→ 3". Otherwise just a hint. The answer text itself is never rewritten.

The grid

Same question. Five hundred times. As a spreadsheet.

Load hundreds of passages — or upload your own list — and ask every row the same question. Each answer has to fit fields you define: text, number, choice.

Illustration of the format

Row (source)Answer, in your schemaCheck
Contract A — § 7 TerminationNotice: 6 months · Special right: yes✓ valid
Contract B — § 4 LiabilityCap: €1m · Exception: intent✓ valid
Contract C — § 9 ExitNotice: — (not found)⚠ field empty

Validated, not trusted

Missing or wrongly typed fields show up right in the cell.

Lists fan out

Five items in one answer become five rows — each still linked to its source passage.

Redo just the bad rows

Re-run single rows with a stronger model; duplicates are merged; export to Excel.

Typical uses: a list of every obligation in a regulation with addressee, deadline and evidence · the same check across hundreds of contracts · first-pass screening of a customer list.

The knowledge graph

A map of who relates to what — with receipts.

The model reads every passage and pulls out things and relations: who is responsible for what, what mitigates which risk, who owns whom. Every single link carries a verbatim quote from the original.

Zoomed in: entities and typed relations such as MITIGATES, OWNS, APPLIES_TO.
Close-up of the Matrix knowledge graph: entities such as risk analysis, human-rights training and water use, connected by relations like MITIGATES and APPLIES_TO.

Quotes are verified

A link whose quote can't be found is repaired first, not silently dropped. If the quote says the opposite ("must not…"), it's flagged.

Duplicates merged

"BaFin" and "Federal Financial Supervisory Authority" become one node — every merge is logged and can be undone by hand.

Lenses, not rewrites

Raw labels stay untouched; a switchable "lens" maps them onto your own vocabulary — no re-extraction needed.

Topics and summaries

The graph is split into clusters, each with a short summary — a table of contents for the whole corpus.

Nothing disappears quietly

Failing passages go to quarantine and are retried; every finding lands in a review queue.

Honest about quality

A local 30B-class model gives a fast first map — orientation, not proof. A frontier model gets close to "audit-ready".

Runs everywhere

One engine. Mac, iPhone, browser.

Everything important lives in a single core written in Rust. The apps on top are thin.

Mac

Local models on Apple's GPU (Metal) and Neural Engine — fast and power-efficient.

iPhone and iPad

The same app, the same features, down to sharing documents into it.

Browser

A Linux container in your own network. Nothing to install for users.

How Matrix answers a questionDocuments go into a phone. On the device they are split, embedded and searched; the answer comes back with the exact source sentence and its location.thousands of pagesairplane modesplitembedsearch & readanswerexact sentence + where it is

Any model

Bring your own AI. Or none from outside at all.

One cockpit for every model connection — each assigned to the job it's good at.

On the device

Search and download models from Hugging Face inside the app; run them locally.

Small specialists

Embedding and re-ranking models for search, also offline.

Company gateway

OpenAI-compatible or Azure endpoints behind your own infrastructure.

Cloud, deliberately

Direct providers where allowed — a conscious, per-case choice.

Per-model limits

Requests per minute, token budgets, parallelism — gentle on gateways and budgets.

Only what it can do

Tool use and reasoning are switched on only for models that actually support them.

Built for bad days

Tested against real outages, not sunny days.

Load balancing

Several endpoints as one pool; throttled members are routed around before a fallback model steps in.

Slow gateway? Stream.

When a proxy gets sluggish, the app switches to streaming so the connection doesn't time out.

Retry with judgement

Temporary errors are retried with growing waits; stubborn cases are skipped instead of killing the run.

Fallback model

If a model returns garbage, a backup model takes over — only for that one passage.

Background jobs

Long runs keep going in the background with visible progress; the app stays usable.

Cool-headed

Local models watch the device's temperature and slow down instead of overheating a laptop.

By the numbers

Measured, not felt.

2 weeksfrom a 48-hour prototype to the working app
65 → 94%top-5 hit rate on a set of legal-norm questions once re-ranking is switched on
~2×faster re-ranking with the GGUF model on Apple silicon, at equal or better quality
4,385entities extracted from a single annual report, each relation backed by a quote

Retrieval quality is tracked with fixed "golden" question sets and expected source passages — quality is judged on those, not on impressions.

Where it stands

Honest status.

Asking and the grid work today, with a local model or a gateway. The knowledge graph is two-tier: local for a fast first map, a strong model where it has to hold up. The checks run either way — they show you where the output can be trusted.

Interested in trying it on your own documents, or in how a part of it works?

Write to me