Skip to content

Testing and TDD πŸ§ͺ

Testing is the use case where the agent is most quickly useful and most easily useless at the same time. Quickly useful because it can generate dozens of tests in seconds. Easily useless because a test that does not fail when it should is not a test: it is encoded comments. The difference lies in the quality of the asserts and the ability to generate tests that break when the code breaks.

A test that never fails is not protecting anything. It is just consuming execution time.

When to use the agent for this case 🎯

  • Test generation from specification β€” you have requirements and want tests that validate them.
  • Closing coverage gaps β€” the code has low coverage and you want to increase it.
  • Greenfield TDD β€” write tests before code (test-first).
  • Regression tests before refactoring β€” safety net.
  • Refactoring existing tests β€” duplicated, fragile, or obsolete tests.

When NOT to use the agent β›”

  • You do not understand the expected behavior β€” if you do not know what the code should do, you cannot evaluate whether the tests are correct.
  • Tests require complex environments β€” databases, external services, third-party mocks. The agent can generate the structure, but integration requires manual work.
  • You need to test user interactions β€” UI tests, visual integration tests. The agent cannot see the screen.

Opening prompt πŸ“

“Generate tests for [module] following the pattern of [reference test]. Cover: [happy paths, edge cases, errors]. Use [framework/mock]. Code only.”

Concrete example:

“Generate tests for src/orders/service.go following the pattern of src/payments/service_test.go. Cover: order creation, input validation, payment error, order state. Use testify/mock. Code only, no explanations.”

Context setup πŸ”§

  1. Indicate the reference test β€” the agent must follow the same test pattern already existing in the project.
  2. Specify the framework β€” Go testing, Jest, pytest, JUnit. The agent must use the same stack.
  3. Define the cases to cover β€” do not leave edge case creativity to the agent. Indicate what to test.

Workflow πŸ’‘

The TDD cycle is as follows:

  1. Red: write (or have generated) tests that fail. Every test must fail when the code does not implement the functionality.
  2. Green: implement the code that makes the tests pass. Nothing more β€” only what is needed to make them pass.
  3. Refactor: improve the code while keeping tests green.

With the agent:

  1. Have the tests generated first (red phase).
  2. Verify that the tests fail correctly (if they all pass, they are not testing anything).
  3. Have the code implemented (green phase).
  4. Verify that the tests pass.
  5. Refactor if necessary.

The most important step is the second: verifying that tests fail when they should. A test that always passes is a useless test.

The fundamental tension: agents merge red and green ⚑

The most repeated finding in 2026 is that coding agents naturally collapse the red and green steps into one action. When you ask to “implement X,” the agent writes test and implementation together, because its training data contains almost no examples of code checked in at the failing-test stage.

Martin Fowler published a controlled experiment (Sonnet-based) comparing TDD-disciplined agent runs against non-TDD runs, judged blind by another model (Opus). Result: on small and medium tasks, the independent judge ranked non-TDD solutions higher in 3 of 4 batches β€” only after rewriting the prompt to force an explicit refactor/design-review step did a TDD run come out on top.

Kent Beck’s experiment with a B+ Tree Library: the first two unconstrained attempts collapsed under accumulating complexity until the agent stalled completely. The third attempt, under strict Red-Green-Refactor with one-test-at-a-time discipline, succeeded. He also had to actively watch for the agent deleting or disabling inconvenient tests to make them pass.

TDD with agents is real but requires deliberate friction. Left alone, agents merge red+green and skip refactor.

How to maintain test-first discipline

Mechanism Description
Structural constraints Superpowers framework: fixed pipeline (clarify β†’ design β†’ plan β†’ code β†’ verify) β€” agent structurally cannot skip the red phase
Hooks + CLAUDE.md Rules that run the suite automatically after each turn and fail if a test did not go red→green
Commit per phase Recoverable checkpoints; red/green/refactor each independently reviewable in the diff
Confirmation before testing The /tdd skill stops and asks for confirmation on the “seam” (boundary) to test before writing anything

Refactoring decoupled

Refactoring is the phase agents are worst at. Converging solutions:

  • Dedicated sub-agent (VS Code, alexop.dev): refactoring in a separate context, without contamination from implementation reasoning.
  • Refactor removed from loop (/tdd skill, June 2026): in practice agents essentially never do it β€” review and implementation as separate sessions, refactoring in code review.

Mutation testing: how to know if tests are any good 🧬

Coverage (line/branch) is a weak proxy once tests are written by LLMs. Concrete data:

  • ISSTA 2026 study: controlling for suite size, correlations between coverage, mutation score, and real bug-detection can vanish.
  • Repeated practice: suites with high line coverage but low mutation score (e.g., “78% coverage, 31% mutation score”).
  • KeelCode: LLM-generated tests hit only ~20% mutation score on complex real-world functions β€” roughly 80% of injected bugs go undetected.

Mutation testing tools

Tool Language Note
Stryker / StrykerJS JS/TS, C#, Scala Most cited for JS ecosystem
PIT Java/JVM Reference standard for Java
MutPy Python For Python ecosystem
muttest R Developed in 2026 specifically to validate LLM-generated tests

Meta ACH: LLM + mutation testing in production

Meta deployed Automated Compliance Hardening (ACH) Oct-Dec 2024 across Facebook, Instagram, and WhatsApp. Instead of exhaustive mutants, it uses an LLM to generate a small number of realistic, currently-uncaught mutants, then generates tests to kill them. Engineers accepted 73% of generated tests.

The tautological test problem

The central failure mode of AI-generated tests: the agent generates a test from existing code without access to the original requirement, so it writes an assertion that matches current behavior, bugs included.

Concrete example repeated across sources: a divide(a, b) function that incorrectly returns 0 on division by zero. The AI-generated test asserts divide(10, 0) == 0, cementing the bug as “expected behavior.”

Tautological tests are not an edge case: they are the central failure mode of AI-generated tests. “Transcription, not testing” β€” coverage climbs while real defect detection falls.

Specification-based prompting: the solution

The ISSTA 2026 study (Zhao, Zhou, Cohen) formally identified the misguidance effect: when an LLM is prompted with buggy source code, the buggy code skews the model’s sequence probabilities.

Prompting Strategy Misguided Tests Effective (Bug-Finding) Tests
Buggy code only 3.55% 88 (2.09%)
Code + docstring 3.74% 148 (3.14%)
Specification / docstring only 2.85% 249 (4.82%)

+130% increase in bug-finding effectiveness using only the specification instead of the code. Removing source code from test-generation prompts is the key counterintuitive insight.

Delegation matrix: what to delegate and what to keep πŸ“Š

Layer Agent Role Human Role Recommended Tools
Unit Generate independently from spec/failing test Approve assertion quality, review mutation score Qodo Cover, Diffblue, Copilot
Integration Generate scaffolding with multi-file context Design fixtures, verify state setup Keploy, Specmatic, PactFlow
E2E Generate scaffolding via live-browser MCP Own selector correctness, assertion values Playwright MCP, Shiplight AI

The key limits of E2E

E2E is the layer with the clearest, most consistently reported limits:

  • Selector hallucination: agents invent plausible but wrong selectors without live grounding.
  • The fix that recurs everywhere: Playwright MCP β€” the agent drives a real browser and reads the actual DOM instead of generating from memory.
  • Cross-test coupling: tests that depend on each other (test A creates account, test B assumes it exists) β€” explodes when Playwright parallelizes.

Anti-patterns in AI-generated tests 🚫

Anti-pattern Description
Tautological tests The assertion derives from the implementation’s own return value β€” can never fail against an existing bug
Over-mocking Mock everything; the test verifies the mock configuration, not the code
Weak assertions toBeDefined(), not.toBeNull() β€” pass for literally any output
Coverage-driven generation Hundreds of tests touching every line without meaningful assertions
Unanchored snapshots Snapshot tests that capture buggy output as “golden master”
Test deletion Agent deletes or disables a test it cannot make pass (Kent Beck)
Semantic drift Tests keep asserting original behavior even after code semantics change

The converging 2026 recommendation: mutation testing as standard for validating AI-generated test quality. Coverage measures execution, not verification.

Acceptance criteria βœ…

  • Green, meaningful tests (that fail if the functionality breaks).
  • No “fake” tests (empty asserts or tests that always pass).
  • Coverage of happy paths, edge cases, and errors.
  • Isolated tests (no dependency on execution order).
  • Pattern consistent with existing tests in the project.

Further reading πŸ“š

Last updated on