Matrix
Ask a thousand pages a question. On a phone, offline.
Give it regulations, contracts or annual reports. Ask anything. You get an answer — and the exact sentence it came from, so you can check it in five seconds instead of trusting it.
The problem
AI sounds right. That's the problem.
Chatbots are fluent and confident — and can't tell you where it says that. For anyone whose job is to be right on paper, "sounds plausible" is worth nothing.
Regulations, contracts and reports are thousands of pages of scattered knowledge where every statement needs a source. Matrix is built for exactly that kind of reading.
The one rule
The AI reads. Everything you could check is plain code.
The language model is only used where language understanding is needed. Splitting documents, following references, verifying quotes, correcting sources and logging every step is done by ordinary rules — reproducible, no AI judgement involved.
Throughout this page: AI means the model decides,rules means deterministic code does.
Every answer shows its sources
Down to the verbatim quote in the original. Clickable, checkable.
Your data stays home
Runs with local models on the device or inside your own network. Nothing has to leave.
Nothing happens in the dark
Every search, every hit, every automatic correction is logged and can be replayed.
Ask a question
It doesn't guess. It goes looking.
Ask once, and the model researches the documents itself — like a junior analyst with a strict supervisor.
Prepare the documents rules
Split along their real structure — articles, sections. Each piece gets a "meaning fingerprint" so search works on sense, not just words. References like "pursuant to Art. 92(1)" are linked to their target.
Search, read, search again AI
The model writes its own search queries, opens passages, follows references and keeps digging with the vocabulary it found — until it has enough evidence or runs out of budget.
Guard rails rules
A search budget caps the cost. Watchdogs stop it from going in circles. An answer given without searching at all is thrown away.
Answer only from what it found AI
Every statement is tagged with its source number: [1], [2], [3].
Check the answer rules
The checks in the next section run on every answer, then everything is shown with clickable sources and a full log.
Built-in fact check
It checks its own homework.
After every answer, a chain of plain-code checks goes over each citation.
Does the source exist?
Every [n] is validated against the list of sources actually retrieved.
Is the quote really there?
Quoted words are searched in the original. Not found → a warning sign on the answer.
Does that source actually say it?
Two independent signals — shared key terms (legal references count triple) and closeness in meaning.
Fix only when both agree
Then the marker is pointed to the right source and labelled "→ 3". Otherwise just a hint. The answer text itself is never rewritten.
The grid
Same question. Five hundred times. As a spreadsheet.
Load hundreds of passages — or upload your own list — and ask every row the same question. Each answer has to fit fields you define: text, number, choice.
Illustration of the format
| Row (source) | Answer, in your schema | Check |
|---|---|---|
| Contract A — § 7 Termination | Notice: 6 months · Special right: yes | ✓ valid |
| Contract B — § 4 Liability | Cap: €1m · Exception: intent | ✓ valid |
| Contract C — § 9 Exit | Notice: — (not found) | ⚠ field empty |
Validated, not trusted
Missing or wrongly typed fields show up right in the cell.
Lists fan out
Five items in one answer become five rows — each still linked to its source passage.
Redo just the bad rows
Re-run single rows with a stronger model; duplicates are merged; export to Excel.
Typical uses: a list of every obligation in a regulation with addressee, deadline and evidence · the same check across hundreds of contracts · first-pass screening of a customer list.
The knowledge graph
A map of who relates to what — with receipts.
The model reads every passage and pulls out things and relations: who is responsible for what, what mitigates which risk, who owns whom. Every single link carries a verbatim quote from the original.
Quotes are verified
A link whose quote can't be found is repaired first, not silently dropped. If the quote says the opposite ("must not…"), it's flagged.
Duplicates merged
"BaFin" and "Federal Financial Supervisory Authority" become one node — every merge is logged and can be undone by hand.
Lenses, not rewrites
Raw labels stay untouched; a switchable "lens" maps them onto your own vocabulary — no re-extraction needed.
Topics and summaries
The graph is split into clusters, each with a short summary — a table of contents for the whole corpus.
Nothing disappears quietly
Failing passages go to quarantine and are retried; every finding lands in a review queue.
Honest about quality
A local 30B-class model gives a fast first map — orientation, not proof. A frontier model gets close to "audit-ready".
Runs everywhere
One engine. Mac, iPhone, browser.
Everything important lives in a single core written in Rust. The apps on top are thin.
Mac
Local models on Apple's GPU (Metal) and Neural Engine — fast and power-efficient.
iPhone and iPad
The same app, the same features, down to sharing documents into it.
Browser
A Linux container in your own network. Nothing to install for users.
Any model
Bring your own AI. Or none from outside at all.
One cockpit for every model connection — each assigned to the job it's good at.
On the device
Search and download models from Hugging Face inside the app; run them locally.
Small specialists
Embedding and re-ranking models for search, also offline.
Company gateway
OpenAI-compatible or Azure endpoints behind your own infrastructure.
Cloud, deliberately
Direct providers where allowed — a conscious, per-case choice.
Per-model limits
Requests per minute, token budgets, parallelism — gentle on gateways and budgets.
Only what it can do
Tool use and reasoning are switched on only for models that actually support them.
Built for bad days
Tested against real outages, not sunny days.
Load balancing
Several endpoints as one pool; throttled members are routed around before a fallback model steps in.
Slow gateway? Stream.
When a proxy gets sluggish, the app switches to streaming so the connection doesn't time out.
Retry with judgement
Temporary errors are retried with growing waits; stubborn cases are skipped instead of killing the run.
Fallback model
If a model returns garbage, a backup model takes over — only for that one passage.
Background jobs
Long runs keep going in the background with visible progress; the app stays usable.
Cool-headed
Local models watch the device's temperature and slow down instead of overheating a laptop.
By the numbers
Measured, not felt.
Retrieval quality is tracked with fixed "golden" question sets and expected source passages — quality is judged on those, not on impressions.
Where it stands
Honest status.
Asking and the grid work today, with a local model or a gateway. The knowledge graph is two-tier: local for a fast first map, a strong model where it has to hold up. The checks run either way — they show you where the output can be trusted.
Interested in trying it on your own documents, or in how a part of it works?
Write to me