Four green checks
Over about a week, preparing a release, this project found four separate controls that were reporting success while being wrong. Not flaky. Not intermittently misconfigured. Confidently, repeatably wrong, each with a passing check on top.
- A production build exited
0having never contacted a database. It assigned a placeholder connection string before reading any environment file, generated every page from static fallbacks, and reported success. Two such builds were taken as evidence the artifact was sound. - The written convention was itself wrong. It prescribed gating build-time warnings on "is this production", which a production build also satisfies. The code implemented that faithfully across twenty-three call sites and produced several hundred stack traces per build.
- An infrastructure inventory documented two protections that did not exist.
- That inventory's own tests enforced the fiction, asserting the name of a branch that had never been created.
The fourth is the one worth sitting with. The test was not absent or skipped. It ran, it passed, and what it protected was a belief.
The common shape
None of these were failures to check. All four were checks that could not fail.
The build could not fail on an unreachable database, because nothing in it treated that as an error; its success carried no information. The convention could not be violated, because the condition was always true. The inventory could not drift, because its test asserted the document against itself rather than against the provider.
Which gives a usable test for your own tooling: a green check is only evidence when a red one was possible. For any check you rely on, you should be able to name the change that turns it red. If you cannot, it is not verifying anything: it is reporting that it ran.
How every one was actually found
Not by better tooling. By comparing what a tool claimed against what the system was observed to do: the provider's API rather than the document describing it, the running site rather than the build log.
Two of the week's wrong turns came from reading the tail of a long log and declaring success. The contradicting warning was near the top both times, in language that read as routine: "Setting dummy connection string for offline compilation." Perfectly clear, printed on every silent build, and missed because it described an intended mode. A warning that sounds like a status line is not a warning.
The fifth one, caused on purpose
The other four were found. This one was manufactured, by finally running a database restore rehearsal that had sat on the backlog specifically because nobody had ever done it.
The rehearsal cut production over. Not a failure of the tool: the parameter controlling that behaviour defaults to "yes, promote this", and the documentation says so plainly. It was read, and the default was accepted anyway, which is its own small lesson about how documentation performs under momentum.
No data was lost, and that was verified rather than assumed: identical schema and content fingerprints before and after, and the newest write in the database predated the restore point by thirty-four days. The rollback was the instructive part. Reversing the designation did not move the compute endpoint with it, and the provider refuses both to delete a root branch's endpoint and to add a second one to an occupied branch. The operation was not symmetric, and nothing said so in advance.
The rehearsal was worth more for that than for the number it produced. A recovery procedure nobody has executed is a hypothesis. This one turned out to have a sixty-four second recovery time and a footgun in the second step, and only one of those was discoverable from the documentation.
What actually changed
Not "be more careful". The useful changes were structural:
- Make the build fail on an unreachable data source, so that a green build means something. Distinguish an empty result from a failed one: they are different states, and conflating them buys false alarms or false confidence.
- Gate on build versus runtime, not production versus non-production.
- Verify controls against the provider, not against the artifact describing the provider.
- For every guardrail, confirm the test fails when the guardrail is removed. Do this by removing it.
- Write down the failure modes that are silent, separately from the ones that crash. They need different detection, and only the second kind gets found by watching for errors.
The uncomfortable part
Every one of these controls was built deliberately, by someone trying to be careful. The build guard, the convention, the inventory, and the test were each a good-faith attempt to prevent exactly the class of problem they went on to conceal.
That is not an argument against controls. It is an argument for periodically asking of each one: what would it look like if this were broken, and would I be able to tell? For four controls here, the answer was that it would look precisely like it looked every day.