Enzo Markarian

Most self-improving setups are ratchets.

They notice a mistake, append a rule, and never look back. Six months later the instructions are a pile of corrections written against a model that no longer exists — and every one is still being paid for on every request. The library gets bigger and the agent gets worse. This is the system I built to run in both directions.

The pipeline

One command from "I want to build X" to reviewed, committed code. The stages are separate because each is independently useful; running them by hand means remembering four invocations and the one gate that actually matters.

Interrogate one question at a time, until the plan holds Spec the absolutes, and the commands that prove them Decompose ordered slices, riskiest one first The one gate a human confirms, because everything past here spends money unattended The loop, unattended one ticket per fresh session · the build must pass · an independent reviewer must agree every rejection is written down, with its reason, and replayed into the next attempt The retrospective reads that record and turns it into changes — including deleting rules that no longer earn their place feeds back into the earlier stages

The dashed line is the whole point. Every other arrow is ordinary automation. The one going backwards — a finished project changing the tooling that runs the next one — is the part almost nobody builds, because it only pays off across projects and it is very easy to build badly.

Why the backwards arrow is dangerous

A finished project is a labeled record of everything that went wrong: every rejection, with a reason written at the moment of failure by a process with no incentive to be kind. That record is the best training signal in the whole setup, and normally it is thrown away when the project directory is deleted.

Feeding it back is obviously attractive and quietly corrosive. Every observation looks like a lesson. Add a rule for each one and the instructions grow without bound — and instructions are not free. They are read on every single request, they compete for attention with the instructions that matter, and most of them exist to patch a weakness in a model that has since been replaced.

So the interesting design problem is not "how do I learn from failure". It is how do I stop learning the wrong things, and how do I remove what I learned when it stops being true. Three rules, and they are the entire system:

01  No change without a receipt

A receipt is a specific thing that actually happened — a failed check, a correction I typed, a rejected review, a change that got reverted. Not "it felt clumsy." And one receipt is not enough: one sighting is an incident, and only a second sighting makes it a pattern.

That rule used to depend on someone remembering the first sighting. It doesn't now: observations are filed as items and matched against existing ones, so the count is a fact the system can show rather than a judgement someone makes. The inbox is also expected to empty — every item ends promoted, dropped with a reason, or waiting with a count. An item that waits long enough without a second sighting gets retired as the one-off it is.

02  No instruction without a guard

A guard is a test case that fails without the change and passes with it. If I can't write one, I don't understand the problem well enough to fix it yet — and saying so is a better outcome than adding a plausible rule.

The check is: run the case with the instruction and without it. If the answer doesn't change, the instruction isn't doing anything, and keeping it costs something on every future request. That test now runs as a harness rather than by hand — cases in both conditions in parallel, graded by a reviewer that never learns which arm it is grading.

03  The audit is allowed to delete

Adding to a set of instructions feels productive. Removing feels risky and shows up as a smaller diff. That asymmetry is exactly why these systems rot, and the only cure is making removal as evidence-driven as addition.

Does it work?

Two numbers, both from real runs rather than from the design document.

4 of 6

Instructions deleted by the audit — out of the six that the same pass had just written. They failed the "does the answer change without it" test. One tool ended the exercise a single line longer than it started.

100% → 60%

A rule that did earn its place: removing it dropped a graded outcome from every trial to three in five, across five trials per condition. Kept on evidence rather than on conviction.

The first number is the one I'd lead with. A system that only ever grows is easy to build and tells you nothing about whether it works. A pass that writes six rules and then deletes four of its own is doing the thing that's actually hard.

Two failures the system had, and what changed

The record was quietly rewritten

Every change is logged with its evidence, and that log is meant to be append-only — because a system that can revise the history of its own changes can erase the evidence that it got worse.

Checking that assumption against a backup showed one entry had been edited in place forty-two seconds after it was written, flipping the record of what happened from "added" to "deleted". The correction was substantively right. The problem was that nothing had stopped it, and nothing would have shown it.

What changed: entries now commit to the one before them, so an edit to any past entry breaks a chain that gets checked. That does not prevent a rewrite — it makes one impossible to hide, which is the property a record of getting worse actually needs.

The root cause was smaller and more embarrassing: the tool defaulted to logging every change as an addition, so a deletion got recorded as its opposite. That default is gone.

The test that decides what survives was measuring nothing

The with-and-without test has an uncomfortable property: the result that licenses a deletion is "no difference" — and "no difference" is also exactly what a broken test produces. A harness that has quietly stopped separating its two conditions reports no difference on everything, and recommends deleting the entire library.

It had stopped separating them. A condition that was supposed to run without the instructions could still read them, so both sides were reading the same file and the clean-looking table meant nothing.

What changed: the harness now carries a case designed to be unanswerable without the instructions, and checks whether the condition that is supposed to lack them answers it anyway. It also checks whether an answer quotes the exact lines it was supposed to be missing. If either fires, it prints no result at all rather than a result that can't be trusted.

It also refuses to give a verdict on too few trials — because at two trials, "no difference" and a coin flip are the same table, and that is the direction that deletes things.

What this is, on a resume

I did not write the harness code, and I've been careful not to imply otherwise anywhere on this site. What I did was define the problem and the constraints: that the system must be able to shrink, that no rule enters without evidence that it changes an outcome, that the record of its own changes must be tamper-evident, and that collection can be automated but judgement cannot — nothing modifies the tooling unattended.

The general skill is the one running through the Wine project and the macOS port too: directing technical work I could not personally perform, while keeping enough judgement over it to notice when the answer is wrong. Here that meant noticing that a system built to improve itself will, left alone, get worse in a way that looks like progress.