A green build is not evidence
A passing test suite tells you that the checks you wrote agreed with the code you wrote. That is a weaker statement than most teams treat it as.
By Fidelio
The most expensive failures in software are not the ones that go red. They are the ones that go green while being wrong, because a green signal ends the investigation. Nobody debugs a passing build.
This is worth stating plainly because the industry has built an enormous amount of tooling on the assumption that a passing check is evidence of correctness. It is evidence of something considerably narrower: that the assertions you thought to write agreed with the code you wrote to satisfy them.
The two ways a check can lie
A check that fails when it should pass is annoying and self-correcting; someone investigates within the hour. A check that passes when it should fail is silent, and silence is indistinguishable from health.
The second kind has a specific shape. The test runs, exercises something, asserts something, and reports success — but the thing it exercised was not the thing under test. A fixture was stale. A mock absorbed the call. The binary being served was built before the edit. The assertion checked that an error was thrown without checking which error.
- The test hit a cached or previously-built artifact, not the current code
- A mock or stub satisfied the call, so the real path never ran
- The assertion was true regardless of the behaviour being tested
- The test asserted that something failed, but not that it failed for the stated reason
- The environment differed from the one the assertion assumed
If a passing run and a broken system would look identical to you, the run is not evidence. It is decoration.
Mutation testing as the cheapest available answer
There is a simple, unglamorous technique that catches most of this: deliberately break the code and confirm the check notices. Delete the authorization redirect and run the auth tests. Remove the logging call and run the logging guard. Invert the condition and see if anything goes red.
A check that survives its own mutation is not protecting anything. It is worth knowing this before an incident rather than during one, and the cost of finding out is a few minutes per check.
The uncomfortable part is how often a suite that has been green for months fails this. The comfortable part is that once you have found a surviving mutant, the fix is usually obvious and small.
What this means for agent-built software
All of the above gets worse when the code and the tests are produced by the same process. An agent that writes both the implementation and the assertions has, without any intent to deceive, optimised for agreement between them. The tests encode the same misunderstanding as the code, and they agree enthusiastically.
This is the argument for separating the reviewer from the builder — not as a process nicety, but because self-assessment cannot detect a shared assumption. A reviewer that did not write the code has no investment in believing the code is correct.
It is also the argument for the deploy gate. A verdict, however good, is an input to a decision. The decision belongs to someone who is accountable for the outcome.