SSerguey Asael Shinder
Java coding notes: the JVM, and writing software that lasts

Serguey Asael Shinder: Done is a claim; the state it left behind is the evidence

· by Serguey Asael Shinder / Serguey Shinder

Microsoft published a benchmark for AI agents this week, ThinkingBox, and its central number is not about AI at all. Across about 121,000 trials, two thirds of the attempts that got the task wrong still finished cleanly, called a tool that changes data, and reported no error. The agent said it was done. The only thing that disagreed was the database.

That gap is older than agents. Every system that reports on its own work has it.

A success message describes the attempt, not the outcome. A batch job prints "28 saved, 0 failed". A deploy step goes green. A migration logs "complete". Each of those lines is written by the code that did the work, about the work, at the moment it believed it had finished. None of them is a reading of the thing the work was supposed to change. When the two match, which is most of the time, the difference is invisible, and that is exactly why it hurts when they do not.

Serguey Asael Shinder: Done is a claim; the state it left behind is the evidence
Done is a claim; the state it left behind is the evidence — Serguey Asael Shinder

The failure that matters is the quiet one. A crash is easy: something throws, a pipeline goes red, someone looks. The expensive failures are the ones that end normally. Files written under the same name overwrite each other and the counter still says twenty-eight. A status is set to "solved" when it should have been "on hold". The run is green and the page it published answers with the old content. Nothing in the success path can tell you about these, because the success path is the thing that went wrong.

So check the state, not the report. For each piece of work, write down the end state it must produce, in terms of the thing it changes: rows in a table, files on disk with distinct contents, a live address that returns the new version. Then read that state back, from outside the code that did the work, and compare. ThinkingBox does this by grading the backend after each run. The same discipline in ordinary code is a test that queries the database after the service call, a job that compares its own count with the files on disk, a deploy that fetches the live page with a cache-buster.

And do it more than once. The benchmark's second lesson is that one success proves the task is possible, not that it is reliable. Some models solved hundreds of tasks at least once and fewer than a hundred every time in twenty runs. Code that touches real records deserves the same suspicion: run the check on every execution, not once when the feature ships.

"Done" is a useful word. It just belongs at the end of the check, not in place of it.