Testing and TDD π§ͺ
Testing is the use case where the agent is most quickly useful and most easily useless at the same time. Quickly useful because it can generate dozens of tests in seconds. Easily useless because a test that does not fail when it should is not a test: it is encoded comments. The difference lies in the quality of the asserts and the ability to generate tests that break when the code breaks.
A test that never fails is not protecting anything. It is just consuming execution time.
When to use the agent for this case π―
- Test generation from specification β you have requirements and want tests that validate them.
- Closing coverage gaps β the code has low coverage and you want to increase it.
- Greenfield TDD β write tests before code (test-first).
- Regression tests before refactoring β safety net.
- Refactoring existing tests β duplicated, fragile, or obsolete tests.
When NOT to use the agent β
- You do not understand the expected behavior β if you do not know what the code should do, you cannot evaluate whether the tests are correct.
- Tests require complex environments β databases, external services, third-party mocks. The agent can generate the structure, but integration requires manual work.
- You need to test user interactions β UI tests, visual integration tests. The agent cannot see the screen.
Opening prompt π
“Generate tests for [module] following the pattern of [reference test]. Cover: [happy paths, edge cases, errors]. Use [framework/mock]. Code only.”
Concrete example:
“Generate tests for
src/orders/service.gofollowing the pattern ofsrc/payments/service_test.go. Cover: order creation, input validation, payment error, order state. Use testify/mock. Code only, no explanations.”
Context setup π§
- Indicate the reference test β the agent must follow the same test pattern already existing in the project.
- Specify the framework β Go testing, Jest, pytest, JUnit. The agent must use the same stack.
- Define the cases to cover β do not leave edge case creativity to the agent. Indicate what to test.
Workflow π‘
The TDD cycle is as follows:
- Red: write (or have generated) tests that fail. Every test must fail when the code does not implement the functionality.
- Green: implement the code that makes the tests pass. Nothing more β only what is needed to make them pass.
- Refactor: improve the code while keeping tests green.
With the agent:
- Have the tests generated first (red phase).
- Verify that the tests fail correctly (if they all pass, they are not testing anything).
- Have the code implemented (green phase).
- Verify that the tests pass.
- Refactor if necessary.
The most important step is the second: verifying that tests fail when they should. A test that always passes is a useless test.
The fundamental tension: agents merge red and green β‘
The most repeated finding in 2026 is that coding agents naturally collapse the red and green steps into one action. When you ask to “implement X,” the agent writes test and implementation together, because its training data contains almost no examples of code checked in at the failing-test stage.
Martin Fowler published a controlled experiment (Sonnet-based) comparing TDD-disciplined agent runs against non-TDD runs, judged blind by another model (Opus). Result: on small and medium tasks, the independent judge ranked non-TDD solutions higher in 3 of 4 batches β only after rewriting the prompt to force an explicit refactor/design-review step did a TDD run come out on top.
Kent Beck’s experiment with a B+ Tree Library: the first two unconstrained attempts collapsed under accumulating complexity until the agent stalled completely. The third attempt, under strict Red-Green-Refactor with one-test-at-a-time discipline, succeeded. He also had to actively watch for the agent deleting or disabling inconvenient tests to make them pass.
TDD with agents is real but requires deliberate friction. Left alone, agents merge red+green and skip refactor.
How to maintain test-first discipline
| Mechanism | Description |
|---|---|
| Structural constraints | Superpowers framework: fixed pipeline (clarify β design β plan β code β verify) β agent structurally cannot skip the red phase |
| Hooks + CLAUDE.md | Rules that run the suite automatically after each turn and fail if a test did not go redβgreen |
| Commit per phase | Recoverable checkpoints; red/green/refactor each independently reviewable in the diff |
| Confirmation before testing | The /tdd skill stops and asks for confirmation on the “seam” (boundary) to test before writing anything |
Refactoring decoupled
Refactoring is the phase agents are worst at. Converging solutions:
- Dedicated sub-agent (VS Code, alexop.dev): refactoring in a separate context, without contamination from implementation reasoning.
- Refactor removed from loop (/tdd skill, June 2026): in practice agents essentially never do it β review and implementation as separate sessions, refactoring in code review.
Mutation testing: how to know if tests are any good π§¬
Coverage (line/branch) is a weak proxy once tests are written by LLMs. Concrete data:
- ISSTA 2026 study: controlling for suite size, correlations between coverage, mutation score, and real bug-detection can vanish.
- Repeated practice: suites with high line coverage but low mutation score (e.g., “78% coverage, 31% mutation score”).
- KeelCode: LLM-generated tests hit only ~20% mutation score on complex real-world functions β roughly 80% of injected bugs go undetected.
Mutation testing tools
| Tool | Language | Note |
|---|---|---|
| Stryker / StrykerJS | JS/TS, C#, Scala | Most cited for JS ecosystem |
| PIT | Java/JVM | Reference standard for Java |
| MutPy | Python | For Python ecosystem |
| muttest | R | Developed in 2026 specifically to validate LLM-generated tests |
Meta ACH: LLM + mutation testing in production
Meta deployed Automated Compliance Hardening (ACH) Oct-Dec 2024 across Facebook, Instagram, and WhatsApp. Instead of exhaustive mutants, it uses an LLM to generate a small number of realistic, currently-uncaught mutants, then generates tests to kill them. Engineers accepted 73% of generated tests.
The tautological test problem
The central failure mode of AI-generated tests: the agent generates a test from existing code without access to the original requirement, so it writes an assertion that matches current behavior, bugs included.
Concrete example repeated across sources: a divide(a, b) function that incorrectly returns 0 on division by zero. The AI-generated test asserts divide(10, 0) == 0, cementing the bug as “expected behavior.”
Tautological tests are not an edge case: they are the central failure mode of AI-generated tests. “Transcription, not testing” β coverage climbs while real defect detection falls.
Specification-based prompting: the solution
The ISSTA 2026 study (Zhao, Zhou, Cohen) formally identified the misguidance effect: when an LLM is prompted with buggy source code, the buggy code skews the model’s sequence probabilities.
| Prompting Strategy | Misguided Tests | Effective (Bug-Finding) Tests |
|---|---|---|
| Buggy code only | 3.55% | 88 (2.09%) |
| Code + docstring | 3.74% | 148 (3.14%) |
| Specification / docstring only | 2.85% | 249 (4.82%) |
+130% increase in bug-finding effectiveness using only the specification instead of the code. Removing source code from test-generation prompts is the key counterintuitive insight.
Delegation matrix: what to delegate and what to keep π
| Layer | Agent Role | Human Role | Recommended Tools |
|---|---|---|---|
| Unit | Generate independently from spec/failing test | Approve assertion quality, review mutation score | Qodo Cover, Diffblue, Copilot |
| Integration | Generate scaffolding with multi-file context | Design fixtures, verify state setup | Keploy, Specmatic, PactFlow |
| E2E | Generate scaffolding via live-browser MCP | Own selector correctness, assertion values | Playwright MCP, Shiplight AI |
The key limits of E2E
E2E is the layer with the clearest, most consistently reported limits:
- Selector hallucination: agents invent plausible but wrong selectors without live grounding.
- The fix that recurs everywhere: Playwright MCP β the agent drives a real browser and reads the actual DOM instead of generating from memory.
- Cross-test coupling: tests that depend on each other (test A creates account, test B assumes it exists) β explodes when Playwright parallelizes.
Anti-patterns in AI-generated tests π«
| Anti-pattern | Description |
|---|---|
| Tautological tests | The assertion derives from the implementation’s own return value β can never fail against an existing bug |
| Over-mocking | Mock everything; the test verifies the mock configuration, not the code |
| Weak assertions | toBeDefined(), not.toBeNull() β pass for literally any output |
| Coverage-driven generation | Hundreds of tests touching every line without meaningful assertions |
| Unanchored snapshots | Snapshot tests that capture buggy output as “golden master” |
| Test deletion | Agent deletes or disables a test it cannot make pass (Kent Beck) |
| Semantic drift | Tests keep asserting original behavior even after code semantics change |
The converging 2026 recommendation: mutation testing as standard for validating AI-generated test quality. Coverage measures execution, not verification.
Acceptance criteria β
- Green, meaningful tests (that fail if the functionality breaks).
- No “fake” tests (empty asserts or tests that always pass).
- Coverage of happy paths, edge cases, and errors.
- Isolated tests (no dependency on execution order).
- Pattern consistent with existing tests in the project.
Further reading π
- Writing prompts that work β how to structure prompts for precise output.
- Greenfield from specification β TDD as a pillar of greenfield.
- Refactoring legacy β tests as a safety net before refactoring.