Make a coding agent write the failing test before the implementation

General Published: Updated:

The four-step test-first workflow for Claude Code, Codex and Gemini CLI, the prompts for each step, what to put in your context file, and a hook that blocks edits to test files.

Verified on Oct 10, 2026 These tools change quickly. Please also check the latest official documentation.
Contents
  1. The four steps
  2. The prompts
    1. Step 1: tests only
    2. Steps 2–3: implement
    3. Step 4: tidy up
  3. Put it in the context file
  4. Claude Code: enforce it with a hook
  5. Codex and Gemini CLI
  6. When it goes wrong
  7. Summary

Ask a coding agent to "implement this feature" and the implementation arrives fast, but the verification often ends at "it should work." The most effective fix is to make it write a failing test first, confirm the failure, and only then implement.

An agent with a concrete goal — "make this test pass" — can iterate on its own. This article covers the prompts that work across all three tools, and the tool-specific ways to enforce the rule rather than request it.

KEY POINT

What you will learn

  • The four steps of the workflow and the prompt for each
  • What to put in your context file (CLAUDE.md / AGENTS.md / GEMINI.md)
  • How to stop the agent from editing the tests, and what to do when it goes wrong

The four steps

StepWhat happensWhat to tell the agent
1. Write the testA failing test derived from the specSay "don't write the implementation yet"
2. Confirm the failureRun it and see it failHave it confirm the failure is "not implemented"
3. ImplementChange the implementation until the test passesSay "don't change the test"
4. Tidy upRemove duplication, improve names, add casesHave it confirm tests still pass afterwards

The important part is that steps 1 and 3 are separate instructions. Roll them into one ("write tests and implement it") and the agent tends to write the implementation first and back-fill a test that fits it.

The prompts

Step 1: tests only

I want a calculateTotal function in `src/lib/pricing.ts` that returns the discounted
total from a quantity and a unit price.
Spec:
- 5% off at 10 or more, 10% off at 100 or more
- Throw if the quantity is 0 or less
- Round fractions down

First, write only the tests for this spec in `src/lib/pricing.test.ts`.
Do not write the implementation yet. Once the tests are written, run
`pnpm test pricing`, confirm they all fail because the function does not exist,
and report back.

Steps 2–3: implement

Implement calculateTotal so every test in pricing.test.ts passes.
Do not modify the test file. If you think a test is wrong, report why
without changing it.
Run `pnpm test pricing` afterwards and paste the result.

Step 4: tidy up

Review the implementation and the tests, and clean up duplication or unclear names.
Add boundary cases if any are missing (9, 10, 99, 100).
Finally run `pnpm test pricing` and `pnpm lint` and report the results.

用語解説

Boundary cases: tests at the points where the spec changes behavior (10, 100) and just either side of them (9, 11). Agents tend to test "representative" values only, so asking for boundaries explicitly pays off.

Put it in the context file

If repeating the instructions gets tiring, write the workflow into your context file. The same text works for all three tools:

## How to implement (test first)
1. For a new feature or a fix, write a failing test first and run it to confirm the failure
2. Implement until the test passes. Do not change the test file while implementing
3. If you judge a test to be wrong, report why without changing it and wait for instructions
4. Include the test command you ran and its result (passed / failed counts) in your report
5. If there is no test setup, propose setting one up first

For keeping one copy of that text across the three files, see CLAUDE.md, AGENTS.md and GEMINI.md differ only in how they load.

Claude Code: enforce it with a hook

In Claude Code you can make this a mechanism rather than a request. The documentation is explicit that a context file is "context, not enforced configuration" and points at a PreToolUse hook to block an action regardless of what Claude decides.

A hook that blocks edits to test files during the implementation phase:

#!/usr/bin/env bash
# .claude/hooks/protect-tests.sh
# Only blocks test-file edits when PROTECT_TESTS=1
[ "${PROTECT_TESTS:-0}" = "1" ] || exit 0
file=$(jq -r '.tool_input.file_path // empty')
case "$file" in
  *.test.*|*.spec.*|*/__tests__/*)
    echo "Test files can't be changed during the implementation phase: $file. If a test is wrong, report why." >&2
    exit 2 ;;
esac
exit 0
{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Edit|Write|NotebookEdit",
        "hooks": [{ "type": "command", "command": "bash .claude/hooks/protect-tests.sh" }]
      }
    ]
  }
}

Exit code 2 on PreToolUse blocks the tool call, so starting with PROTECT_TESTS=1 claude makes tampering with the tests physically impossible rather than merely discouraged. For hook basics, see Run lint and format on save with Claude Code hooks.

Fix the procedure in a skill: writing the four steps into a /tdd <spec> slash command keeps every run on the same track (Create your own slash commands). Delegating the test run to a subagent keeps long test output out of your main context.

Codex and Gemini CLI

Neither has an equivalent of hooks, so you run this on instructions plus inspection.

  • Codex: keep the implementation phase on --sandbox workspace-write, and check with /diff that the test files are unchanged before you approve. Put the workflow in AGENTS.md.
  • Gemini CLI: turn checkpointing on, and if a test file was modified, roll back with /restore. Put the workflow in GEMINI.md.

When it goes wrong

SymptomWhat to do
It writes the implementation during the test stepRepeat "don't write the implementation yet" at the end of the prompt too. In Claude Code, use plan mode to get a plan only
It loosens the test when it won't passState "don't change the test" and block it with a hook
The tests are all mocks and prove nothingSay "mock external dependencies if you must, but never mock the logic of the function under test"
The test run takes too longPut a narrowing command in the prompt (pnpm test pricing)

Summary

  • Make "write the test" and "implement" separate instructions, with a confirmed failure in between
  • State "don't change the test file" and "include test results in your report" every time, or put them in your context file
  • Claude Code can block test-file edits with a PreToolUse hook and fix the procedure in a skill
  • With Codex and Gemini CLI, use /diff or checkpointing to check the tests weren't tampered with

FAQ

Won't an AI-written test just be shaped to fit the implementation?
Not if you make it write the test first and confirm the test fails before any implementation exists. The order of the instructions is what does the work.
Will the agent rewrite the test to make it pass?
It can. Say "do not change the test file" explicitly, and in Claude Code block edits to test files with a PreToolUse hook.
Does this work in a project with no tests yet?
Yes. Start by asking for a working test setup, get one test running, then begin the loop.

Primary sources

This article was drafted by AI from official documentation and reviewed by the site operator before publishing. Found a mistake? Let us know via the contact page.