The document is not the control
A repository convention exists in one of three states: written down, implemented, or enforced. They are routinely confused, and an agent has no way to tell them apart: it reads the convention, assumes it holds, and builds on it.
A concrete example from this repository. The branch-naming convention was documented. A validateBranchName() function existed. A unit test asserted it worked, and passed continuously. Nothing anywhere called it. The convention was written, implemented, and tested. It was also unenforced for months, while every artifact suggested otherwise.
That is the characteristic agent-first failure, and it is worse for agents than for humans. A human who has been on the team a while knows which rules are real. An agent has only the artifacts, and the artifacts said it was enforced.
Test that the control is invoked, not that it works
This inverts the usual advice about testing behaviour over implementation, and the inversion is the point.
A test asserting validateBranchName("feat/x") returns true tells you the logic is right. It tells you nothing about whether anything calls it. For a guardrail, "is it wired in" is the property that matters, because unwired logic fails open and silently.
// Weak: passes forever while the hook does not exist.
expect(validateBranchName("nonsense")).toBe(false);
// Strong: fails the moment the guardrail is disconnected.
const hook = readFileSync(".husky/pre-push", "utf8");
expect(hook).toMatch(/validate-branch/);
Source-level assertions feel crude. They are appropriate here precisely because they assert a fact about the system rather than about a function; the fact is the thing that keeps being untrue.
Make the remedy executable
When a gate fails, the message it prints is the entire interface. An agent will do exactly what it says, so a message that describes a problem without naming a command produces a guess.
"Documentation is out of sync" sends an agent looking for the files. "Run npm run doctor:fix" gets it fixed in one step. The remediation command usually already exists; it is just not named at the point of failure. The cost of that omission compounds: every failure spends a full cycle rediscovering it.
Scope selection is a correctness property
Running only tests related to changed files is a sensible optimisation, and it has a sharp edge: relatedness is computed through the import graph. A change to a build script that nothing imports has no related tests, so the pre-commit hook runs none and reports success.
In this repository that precise gap shipped a regression that printed a live production credential into test output. The hook passed. It was correct to pass, given its rules. The rules were wrong.
The fix is not to abandon scoped testing. It is to recognise that some paths (build scripts, configuration, seeds, and anything widely depended upon but imported by nothing) must force a broader run regardless of what the graph says.
A short checklist
- For every documented convention, name the file and line that enforces it. If you cannot, it is documentation, not a control.
- For every guardrail, write a test that fails when it is disconnected, then verify the test fails by disconnecting it.
- For every failure message, include the exact command that fixes it.
- For every scoped optimisation, enumerate what it deliberately skips and decide whether that set is acceptable.
- Prefer a loud failure to a quiet fallback anywhere an agent will read the outcome as success.
None of this is specific to agents. Agents just remove the slack that made the gaps survivable: they do not know which rules are real, so unenforced conventions get discovered at the worst possible moment.
PromptOps covers versioning and evaluating prompts as build artifacts; Hono Kiln covers the monorepo boundaries that make scoped checks tractable in the first place.