Tests are the only fixed point left when the implementation is disposable
When an AI can rewrite the implementation at will, where does the spec live? Grounded in the docs that say instructions carry no enforcement, and in what a hook can block.
Contents
The more disposable the implementation becomes, the more the spec moves into the tests. That is the claim.
Working with an AI coding tool, the same feature gets rewritten several times in a day. What gets rewritten is not the spec. The spec is whatever survives and can tell you that you are wrong: tests, types, and the checks you put in hooks.
Everything below comes from the Claude Code documentation as fetched on 2026-10-09.
KEY POINT
What you will learn
- Why an instruction cannot be a fixed point, in the docs' own words
- The line between advisory and fails-the-build
- When this claim does not hold
Instructions are not a spec
A promise written in CLAUDE.md is not a spec. The memory docs say of CLAUDE.md and auto memory: "Claude treats them as context, not enforced configuration," and add: "To block an action regardless of what Claude decides, use a PreToolUse hook instead."
The permissions page is blunter: "Permission rules are enforced by Claude Code, not by the model. Instructions in your prompt or CLAUDE.md shape what Claude tries to do, but they don't change what Claude Code allows."
So "please build it this way" is a statement of intent. It might be read, and it might be followed.
Only what a machine can fail is a fixed point
What is a fixed point, then? Whatever can fail. The per-event hook behavior makes the line explicit.
| Event | On exit code 2 | Usable as a fixed point |
|---|---|---|
PreToolUse | Blocks the tool call | Yes — stops it before it runs |
PostToolUse | The tool already ran; stderr is shown to Claude | Partly — can't stop it, can make Claude fix it |
PostToolBatch | Stops the agentic loop before the next model call | Yes |
Stop | Prevents Claude from stopping, continues the conversation | Yes — "never finish while red" |
InstructionsLoaded | Exit code is ignored | No |
The one that pairs best with a test suite is Stop: "Prevents Claude from stopping, continues the conversation." A hook that runs the suite and exits 2 when it is red turns "don't report done while tests fail" from a request into a mechanism. That gap is the asymmetry above.
PostToolUse cannot stop anything — "Shows stderr to Claude; the tool already ran." The edit happens; exit 2 shows stderr so Claude fixes it. Linters and formatters belong there; see Run lint and format on save with Claude Code hooks. In runs nobody watches, this is the only defense you have (Run Claude Code headless in CI).
The docs themselves ask for verification targets
The same conclusion arrives from the cost page: "Give verification targets: Include test cases, paste screenshots, or define expected output in your prompt. When Claude can verify its own work, it catches issues before you need to request fixes."
That is not advocacy for test-first development; it is a way to stop wasting tokens. The demand is identical: fix the expected value first. The reasoning matches the stated case for plan mode, "preventing expensive re-work when the initial direction is wrong" (Claude Code plan mode).
This site runs on the same structure: every article must pass a lint script and a build before publication, and length, heading counts and internal link targets fail there. A promise you want kept has to be shaped so it can fail.
Diff review is not a fixed point
A human reading the pull request diff no longer holds the position it used to. Each rewrite replaces the diff wholesale, so it cannot tell you whether last round's review comment survived. Only converting that comment into a test, a type or a hook can.
This is not an argument against review. It is a relocation: turning review output into something that fails becomes part of the review. For the restraining side, see Design permissions in Claude Code's settings.json.
Objections and limits
The claim fails under these conditions.
When the AI writes the tests too. Tests written while looking at the implementation are a tautology of it. The fixed point is only the part a human decided beforehand. Tests generated afterwards work as a regression net, not as the spec of record.
Domains a test cannot express. Visual appearance, how obvious an interaction is, perceived performance, accessibility, security posture. You either build a different fixed point — screenshot comparison, a benchmark threshold, an audit checklist — or a human keeps looking.
When the fixed point costs more than it returns. The first test in a legacy codebase with no test infrastructure is expensive, and a throwaway prototype never recovers it. Not building one is sometimes right — but as a stated choice, not an unnoticed omission.
When you pick the wrong thing to fail on. Hooks carry enforcement; what fails is still your decision. Design around the belief that PostToolUse blocks, and all you get is a warning after the edit landed.
Summary
- The docs call CLAUDE.md context, not enforced configuration, so an instruction is not a fixed point
- Only what a machine can fail counts:
PreToolUseandStopblock,PostToolUseonly makes Claude fix it - Even on cost grounds, the docs ask you to give verification targets up front
- Every rewrite replaces the diff, so review output must become something that fails
- It does not hold when the AI writes the tests, in domains a test cannot express, or when the fixed point does not pay for itself
FAQ
- What if I don't have time to write tests?
- Fix one expected value first. The official documentation says that including test cases or expected output in your prompt lets Claude catch issues before you have to request fixes. You do not need coverage to get the benefit.
- Isn't it enough to put "don't report done while tests fail" in CLAUDE.md?
- No. The documentation calls CLAUDE.md context, not enforced configuration, and points you at a PreToolUse hook to block an action regardless of what Claude decides. To stop the turn from ending, use a Stop hook.
- Could tests become unnecessary in the AI era?
- Unlikely. The cheaper it gets to produce an implementation, the more the value shifts to whatever decides which implementation is correct. But as the limits section says, domains a test cannot express need a different fixed point.
Primary sources
This article was drafted by AI from official documentation and reviewed by the site operator before publishing. Found a mistake? Let us know via the contact page.