Make a coding agent write the failing test before the implementation
The four-step test-first workflow for Claude Code, Codex and Gemini CLI, the prompts for each step, what to put in your context file, and a hook that blocks edits to test files.
Contents
Ask a coding agent to "implement this feature" and the implementation arrives fast, but the verification often ends at "it should work." The most effective fix is to make it write a failing test first, confirm the failure, and only then implement.
An agent with a concrete goal — "make this test pass" — can iterate on its own. This article covers the prompts that work across all three tools, and the tool-specific ways to enforce the rule rather than request it.
KEY POINT
What you will learn
- The four steps of the workflow and the prompt for each
- What to put in your context file (CLAUDE.md / AGENTS.md / GEMINI.md)
- How to stop the agent from editing the tests, and what to do when it goes wrong
The four steps
| Step | What happens | What to tell the agent |
|---|---|---|
| 1. Write the test | A failing test derived from the spec | Say "don't write the implementation yet" |
| 2. Confirm the failure | Run it and see it fail | Have it confirm the failure is "not implemented" |
| 3. Implement | Change the implementation until the test passes | Say "don't change the test" |
| 4. Tidy up | Remove duplication, improve names, add cases | Have it confirm tests still pass afterwards |
The important part is that steps 1 and 3 are separate instructions. Roll them into one ("write tests and implement it") and the agent tends to write the implementation first and back-fill a test that fits it.
The prompts
Step 1: tests only
I want a calculateTotal function in `src/lib/pricing.ts` that returns the discounted
total from a quantity and a unit price.
Spec:
- 5% off at 10 or more, 10% off at 100 or more
- Throw if the quantity is 0 or less
- Round fractions down
First, write only the tests for this spec in `src/lib/pricing.test.ts`.
Do not write the implementation yet. Once the tests are written, run
`pnpm test pricing`, confirm they all fail because the function does not exist,
and report back.
Steps 2–3: implement
Implement calculateTotal so every test in pricing.test.ts passes.
Do not modify the test file. If you think a test is wrong, report why
without changing it.
Run `pnpm test pricing` afterwards and paste the result.
Step 4: tidy up
Review the implementation and the tests, and clean up duplication or unclear names.
Add boundary cases if any are missing (9, 10, 99, 100).
Finally run `pnpm test pricing` and `pnpm lint` and report the results.
用語解説
Boundary cases: tests at the points where the spec changes behavior (10, 100) and just either side of them (9, 11). Agents tend to test "representative" values only, so asking for boundaries explicitly pays off.
Put it in the context file
If repeating the instructions gets tiring, write the workflow into your context file. The same text works for all three tools:
## How to implement (test first)
1. For a new feature or a fix, write a failing test first and run it to confirm the failure
2. Implement until the test passes. Do not change the test file while implementing
3. If you judge a test to be wrong, report why without changing it and wait for instructions
4. Include the test command you ran and its result (passed / failed counts) in your report
5. If there is no test setup, propose setting one up first
For keeping one copy of that text across the three files, see CLAUDE.md, AGENTS.md and GEMINI.md differ only in how they load.
Claude Code: enforce it with a hook
In Claude Code you can make this a mechanism rather than a request. The documentation is explicit that a context file is "context, not enforced configuration" and points at a PreToolUse hook to block an action regardless of what Claude decides.
A hook that blocks edits to test files during the implementation phase:
#!/usr/bin/env bash
# .claude/hooks/protect-tests.sh
# Only blocks test-file edits when PROTECT_TESTS=1
[ "${PROTECT_TESTS:-0}" = "1" ] || exit 0
file=$(jq -r '.tool_input.file_path // empty')
case "$file" in
*.test.*|*.spec.*|*/__tests__/*)
echo "Test files can't be changed during the implementation phase: $file. If a test is wrong, report why." >&2
exit 2 ;;
esac
exit 0
{
"hooks": {
"PreToolUse": [
{
"matcher": "Edit|Write|NotebookEdit",
"hooks": [{ "type": "command", "command": "bash .claude/hooks/protect-tests.sh" }]
}
]
}
}
Exit code 2 on PreToolUse blocks the tool call, so starting with PROTECT_TESTS=1 claude makes tampering with the tests physically impossible rather than merely discouraged. For hook basics, see Run lint and format on save with Claude Code hooks.
Fix the procedure in a skill: writing the four steps into a /tdd <spec> slash command keeps every run on the same track (Create your own slash commands). Delegating the test run to a subagent keeps long test output out of your main context.
Codex and Gemini CLI
Neither has an equivalent of hooks, so you run this on instructions plus inspection.
- Codex: keep the implementation phase on
--sandbox workspace-write, and check with/diffthat the test files are unchanged before you approve. Put the workflow inAGENTS.md. - Gemini CLI: turn checkpointing on, and if a test file was modified, roll back with
/restore. Put the workflow inGEMINI.md.
When it goes wrong
| Symptom | What to do |
|---|---|
| It writes the implementation during the test step | Repeat "don't write the implementation yet" at the end of the prompt too. In Claude Code, use plan mode to get a plan only |
| It loosens the test when it won't pass | State "don't change the test" and block it with a hook |
| The tests are all mocks and prove nothing | Say "mock external dependencies if you must, but never mock the logic of the function under test" |
| The test run takes too long | Put a narrowing command in the prompt (pnpm test pricing) |
Summary
- Make "write the test" and "implement" separate instructions, with a confirmed failure in between
- State "don't change the test file" and "include test results in your report" every time, or put them in your context file
- Claude Code can block test-file edits with a
PreToolUsehook and fix the procedure in a skill - With Codex and Gemini CLI, use
/diffor checkpointing to check the tests weren't tampered with
FAQ
- Won't an AI-written test just be shaped to fit the implementation?
- Not if you make it write the test first and confirm the test fails before any implementation exists. The order of the instructions is what does the work.
- Will the agent rewrite the test to make it pass?
- It can. Say "do not change the test file" explicitly, and in Claude Code block edits to test files with a PreToolUse hook.
- Does this work in a project with no tests yet?
- Yes. Start by asking for a working test setup, get one test running, then begin the loop.
Primary sources
This article was drafted by AI from official documentation and reviewed by the site operator before publishing. Found a mistake? Let us know via the contact page.