Software Quality: When the Green Check Proves Nothing

On 1 September 2026 we changed a button on our own website. Then we measured in the browser whether font size, colour and spacing were right.
They were. The check was clean, and it was worthless anyway.
To measure the button at all, we had first switched off an animation that fades it in as you scroll. That animation was exactly what our change had broken. We had therefore switched off the very thing we needed to see, and then confirmed that everything was fine.
The fault was found by a second agent reading the finished code. Not by our measurement.
Why a green check proves nothing
A check is a claim at heart. It says: if something breaks, I will fire. Green is therefore only information when red was possible at all. If it was not, green says precisely nothing, and it says that nothing very convincingly. Nobody cheated here. The setup that made the measurement possible was the same setup that made the fault invisible. Whoever checks always builds those conditions themselves, and when the conditions switch off the very thing at issue, what you end up checking is your own preparation.
How often that happens in other teams, we do not know. With us it happened twice in a single day.
The second case: thirteen green checks that secured nothing
On the same day we were working on our own tooling, on the script that merges finished changes. Its safety rests on a single ordering. First the current state is produced, then the checks run over it. In the wrong order, code goes through that nothing has checked.
We wrote thirteen new checks for it, and all of them were green. The reviewing agent, however, then moved a single line so that the ordering was wrong and ran that part of the suite again. Green again, 168 checks in that part without a single failure. The thirteen new ones had never touched the ordering.
The test that was missing recreates the earlier state artificially and looks at whether the work really is checked again (it has to reproduce the result of the previous check to do so, which makes it more awkward than the rest and is also what makes it useful, because it can only fire in exactly that state).
Having added it, we destroyed the ordering deliberately. Three checks went red. Then we put the code back.
Two rules since 1 September
Since 1 September 2026 those two cases have been a fixed working rule for us, in two parts.
- A new test has to have been red once. Whoever writes it deliberately breaks what the test is meant to protect, watches it fail, and restores the code.
- The setup of a check must not switch off what the check is about. Otherwise you are measuring your own preparation.
The second part comes from the button, the first from the thirteen checks. Both rules cost about a minute per check. They are nevertheless the difference between a check and the memory of a check.
What this means for your own software
For you as the client this is not a technical nicety, but the question of whether your supplier's assurances are covered. Software quality is happily reported in numbers: this many tests, this much coverage, everything green. None of those numbers says whether a single one of those checks was ever able to raise the alarm.
Four questions will help you settle that in a status meeting without any software knowledge.
- When was this check last red? A check that has never fired is not something anyone should rely on.
- Show me how it fails. A short, deliberate error answers the question in two minutes during the meeting.
- What gets switched off so the check can run? Whatever is named here is unchecked.
- Which part of our workflow is covered by no check at all? An honest answer names a gap, not a number.
If we had to choose, we would take three checks that can demonstrably fail over thirty green ones. The thirty cost more time and say less. This becomes particularly clear when AI writes the code, because tests then appear just as quickly and just as plausibly as the code itself.
Whoever checks must be independent, and the check must be able to fail
Neither fault was found by whoever made it.
That a second reviewer is necessary is one half. How much review matches the risk is a separate judgement and lies outside this article. The other half is rarely said out loud: even an independent check is worthless as long as nobody has shown that it can fail.
That is why both sit side by side in our work.
A second agent, ideally on a different model, reviews the first one's work, a person signs off, and every single check has to have proven once that it can go red.