programmiere

The maintainer loop

My software maintains itself. It still has to ask me.

An AI reads the support forum, writes fixes, drafts replies and prepares releases fora WordPress plugin with about 4,000 installs. Then it stops and texts me. One word — "ja" — and exactly that one thing happens. Nothing else.

How the maintainer loop worksA scout reads forums and reviews, a worker writes a fix on its own branch, deterministic checks verify it, and nothing is released, merged or posted until the owner approves that single action from his phone.Scoutreads the forumsWorkerwrites the fixChecksno model involvedGate„ja“one word per action,from my phonethe model only ever proposes —it never ships anything itself

The question everyone asks

How much power do you give an AI agent?

Enough to do real work. Never enough to do damage on its own.

The answer here is a strict split: the model is allowed to propose anything. Deciding, executing and anything that can't be undone happens in plain code — and the irreversible steps need a human, every single time.

Four roles

A small team, minus the meetings.

The parts talk only through files. No shared memory, nothing hidden.

  1. Scout rules

    Polls the plugin's forum and reviews, competitors' changelogs, form-builder forums and its own live spam data. Files anything new into an inbox. No AI involved.

  2. Worker AI

    A headless coding session that implements one item on its own branch. Then a deterministic verifier — not the AI — decides whether the result is even allowed to be proposed.

  3. Concierge AI + rules

    Talks to me over Signal. The model may phrase things and suggest "go" or "veto" for one existing item — it can never execute anything.

  4. Gate rules

    The only door to the outside world: release, merge, forum post. One approval per action, tied to my number. An approval never carries over to the next thing.

Paranoid by design

Assume the internet is trying to talk to the AI.

Forum posts are written by strangers. Some of them would love to tell an AI what to do.

Untrusted text is boxed

Anything from the outside reaches the model clearly fenced off as untrusted data — never as instructions.

The important sentences aren't AI-written

What I'm actually approving is rendered by fixed templates. No model and no forum post can change the wording of the question I say "ja" to.

One question at a time

A single, short-lived confirmation slot. If a new notification arrives while a question is open, the old question is cancelled — so a "ja" can't land on the wrong thing.

Releases need a different word

Publishing to WordPress.org can't be confirmed with the word that approves a merge. Different risk, different key.

No nagging

A pending release shows up in the daily digest — but there is no escalating reminder, because the pace of changes is partly driven by text from strangers.

Fails closed

Unreadable state, broken model, missing marker: the safe default wins, every time.

From the phone

Shipping a release by text message.

I type "release 5.2.0". The loop prepares everything in an isolated copy: version bump, changelog from the commit history, validation. Then it shows me what will go out — and waits.

The changelog gets one optional AI polish for better English. The raw text is shown right next to it, so I judge a difference instead of trusting a rewrite. If the model doesn't answer, the release still goes out with the raw text — a release is never blocked by a model.

Watching the watcher

"Not measured" is not the same as "fine".

Small canary tests check every day that the AI's sandbox still holds and that the model still behaves. When a check can't run, it says exactly that — it never reports "healthy" for something it didn't test.

Lesson learned the honest way: a loop that reports its own failures correctly still needs a second channel to reach its owner. The interesting engineering in agent systems is often not the agent — it's the monitoring around it.

By the numbers

Not a weekend demo.

~30,000lines of loop code across almost 100 modules
1,889automated tests — the test runner refuses any real model call
4roles: scout, worker, concierge, gate
1 wordbetween a prepared change and the outside world

The takeaway

Give agents real work. Keep the keys.

The pattern is general: the model proposes, deterministic code decides what's even possible, and a human holds the one key that can't be taken back. It works for a WordPress plugin — and for anything where an AI should act, but not alone.

Talk to me about it