The number stays green

A test proves your system still does what you told it. Nothing proves what you told it is still true.

By ·

Every system that fails in production passed its tests the morning it failed.

Engineers say a version of this to each other the way you do after a bad night. It is not really a joke, or even irony — it is closer to a definition. A test is a promise a developer makes about the code, and a green test is that promise kept. When the thing finally breaks, it is rarely because a promise was broken; it is because a promise nobody thought to question was kept perfectly, all the way down.

In March, Amazon's retail site — the store, not AWS — went down in a handful of outages across a single week. The Financial Times reported that AI-written code had caused them. Amazon issued an interesting correction: none involved AI-written code. One of the incidents, they wrote, involved AI-assisted tooling, which related to an engineer following inaccurate advice that an AI tool inferred from an outdated internal wiki. Read it slowly, because nothing in that sentence malfunctioned. An engineer asked. A tool answered. The advice was followed, correctly. And it came from a page that had quietly stopped being true. The code was faithful to the wiki; the wiki was no longer faithful to the world; and there was nothing in the loop built to feel the gap between the two. Then look at the remedy. Amazon says it addressed the issue and updated its internal guidance to keep it from happening again — and it went out of its way to deny it had put any new approval gate on engineers using AI tools. No new gate, no better model: clearer words, for the people who read them. One of the largest engineering organisations on earth, correcting in exactly one direction. Back toward the person.

For decades, software engineering has gotten very good at one half of the job and called it the whole. Everything the field builds to keep software honest — the tests, the schemas, the pipelines, and lately the tools that treat the written specification as the source and regenerate the code to match it — checks one thing: whether the system is still faithful to what you already told it. And here the trouble is older than any tool: we gave both halves the same name. We called that faithfulness correctness. It was coherence — the system agreeing with itself. Correctness was the other thing entirely: the agreement between what you declared and what is actually true out in the world this quarter. One word for two different questions. The declaration and the world begin as one claim and drift into two — the same on Monday, strangers by spring — and nothing sounds an alarm as they part, because every instrument we own is pointed at the half that still holds. There is plenty that watches whether the system is up, whether the numbers moved, whether the pattern looks wrong, whether you can trace where each answer came from; there is nothing that watches whether the thing you declared is still true. Not out of laziness — no instrument can be built to close that gap: you can measure whether the running system still matches what you declared, never whether what you declared still matches the world. Where the instrument should be, there is a person.

I have stood there, long before I had a name for it. I was running five finance departments through a hard year — new systems going in across the company, the pressure climbing — and I reached for one by hand: a tracker, a line for every project, kept up by the people already stretched thinnest. It had the shape of knowing and almost none of the substance. The cells filled and overflowed, the file grew, there was never time to keep it true; it gave me the comfort of thinking nothing was slipping past me and told me very little that actually was. And it cost the people filling it the one thing I had spent years defending: the autonomy of a team good enough not to need watching. The tracker had not failed — it did what a tracker does. The mistake was mine, for asking an instrument of coherence to hand me a truth it was never built to hold. Not a line of it ran on a computer; no machine had anything to do with what went wrong.

None of this is the machinery failing. Coherence is a real achievement, and the instruments do exactly what they were built for — they can only be built for the half that stays still. You can hand the person sharper tools; you can even build one for the half that does have an answer — a gauge of how far the running system has drifted from what it declared, not whether the declaration is still true. We built ours. The limit is honest and deliberate: it catches the failures that are loud — a database that dies and reads zero, a service that stops answering — and it is nearly blind to the quiet ones, the Amazon kind, where nothing breaks and everything runs and the answer has just stopped being true. A system can drift a long way and still read green: the number only twitches when something actually breaks, not when it quietly goes wrong. What the gauge does, it does better than a tired person at seven on a Friday: it never assumes the migration probably ran.

One morning two of our own instruments hit the same blind spot and split on it. One was a plain release dashboard: every light on it green, and not one of them watching whether a single thing had actually shipped — which, that morning, nothing had. The green answered the question it was built for, and we had mistaken it for a bigger one. Failure, it turns out, can wear the exact colour of success. The other instrument was the gauge — the number itself. It hit the same gap: part of what it needed to weigh had dropped out of sight. But instead of going green over the hole, it handed us a number for what it could still see and said, plainly, that it was partly blind. And that is where even the honest one ran out — it can tell you it is half-blind; it cannot tell you whether the half it missed was the half that mattered. Someone had to look at that half-lit green and decide whether to trust it, and there was no one there to do it. Every system that fails in production passed its tests the morning it failed; ours passed them on the morning nothing shipped.

What the machine cannot do is be the person — the one who answers for the call. Coherence is a question engineers can answer; correctness is one only the business can. And it does not come down to being clever. Judgment involves plenty of that — reading the room, weighing what matters, catching the quiet thing in the numbers that changes what they mean — and a machine can be taught a fair imitation of all of it. What it cannot be handed is answerability: being the one the loss lands on when the call turns out wrong. That is what disciplines a judgment and makes it worth trusting, and it does not transfer to something that cannot be hurt by being wrong. It is not a capability you bolt on with a bigger model; it is a position, and only a person can hold it.

I came across Will Larson's writing recently, reading everything I could about the world Ficus had pulled me into, and one line stopped me. After laying out the discipline of trusting nothing you have not checked yourself, he named its price almost in passing: this is an exhausting way to live. I knew that sentence better than I would like to. I did not come to this from engineering — I trained as a philosopher, worked for years in publishing and journalism, and then spent years inside finance, answering for systems whose insides I trusted rather than read. The thread through all of it was older than any of the jobs: the need to know whether the thing you are holding up is actually true. What exhausts is not the dramatic version — the crisis, the all-nighter — it is the daily one: the unending work of holding enough of a business in your head to decide whether the business is still the right one, and deciding anyway on the mornings you cannot be sure you have seen all of it. It wears not because the rigour falls short but because the work has no floor: some of what you carry can be checked, but the deepest part cannot, because there is no outside record of whether a business is still true. You never reach the point of having checked enough. And a demand with no floor, carried by a person, does not merely tire — left alone it wears away the very judgment it means to protect, the way endless practice can drain the music out of the player. The relief is not a machine that decides for you; it is one that will do the tireless, mechanical half without end, so that the part of you that judges is still alive when the judging has to be done.

So the honest machine is not the one that pretends to judgment. It is the one that does the measurable half exactly and is built — on purpose, by construction — to refuse the rest. When we needed a model to weigh something you cannot put in a rule — whether a test really checks the thing it claims to check — we tied its hands with one line: cite the evidence or say nothing. Never guess. It has weight because it is not about manners: anything the model cannot point to — a real place a person could go and check for themselves — it simply is not allowed to claim. That does not make it well-behaved. It makes it possible to catch out, which is an oddly useful thing for a machine to be. Its receipts, in that role, are dull and specific, and we are glad to hand them over. A machine that knows the exact edge of what it knows is a stranger thing than a clever one, and a rarer one, and in production it is worth more. Much of the field is building the opposite — systems that present no edge at all, that answer every question in the same voice whether or not they can see the answer. That confidence is the product, right up until the morning it is the disaster.

The Greeks carved three lines into the temple at Delphi. Two are still quoted — know yourself, and nothing in excess — and a machine can be built to obey both: to know the edge of its own sight and stay inside it, measuring, without overstepping into judgment. That self-knowledge is harder to build than confidence, and it is most of what we spent our time on. The third is the one that gets forgotten, and it is not the same kind. The first two are instructions; the third barely resolves into one — give a pledge and ruin is near; or, just as fairly, avoid certainties. It tells you less what to do than what to fear, and leaves the reading to you. That openness is the point: the single maxim that refuses to harden into a rule is the one warning against treating anything as settled — and a machine cannot obey it, because doing so is nothing but judgment. The ruin did not live in the number. It is in letting the green harden into a pledge, something you lean on and stop questioning, and handing it a call it was not carved to hold. That temptation was ours first: to let the number be the answer instead of the measure.

This only sharpens as we hand the doing to agents. An agent can make the change, run the tests, reconcile the ledger — all of it, faster than you would. What it cannot be is the one who stands behind the work the next morning, who says I authorised that, against everything we knew that day, and it was the right call. The more the machine does, the more that sentence weighs, and the less the machine can be the one to say it. The doing keeps getting cheaper, heading for free. The weight of standing behind it does not move at all: doing is a task, and tasks get cheaper; answering for the result is a position, and a position cannot be automated.

None of this makes the machine smaller. A system that measures itself honestly — that hands you the exact distance from your own contract and then stops, clean, at the line where its sight ends instead of guessing past it — is doing the hardest, most truthful thing software can do. It tells you the truth about itself, including where it stops being able to tell. That is what Ficus builds. The judgment stays where it has always been — with the one who has something to lose when it is wrong.

And here is the turn the green hides. A number you cannot trust is the easy one — you go and check it against the world. The hard one is the number you can: honest, complete, every signal it could weigh weighed, and green. It can hand you that green and it still cannot tell you the drift is fine — that this once the world moved and the system was right to hold its ground. Someone still has to lean back, look at the green number, and decide it is lying — not about what it measured, where it is honest, but about what it now means: whether what drifted was a rule that still holds, or a habit nobody went back to question. That was always the job. It was never going to be the machine's.