
Simon Willison just published a new guide on agentic engineering patterns. The first entry: Red/Green TDD. The core claim is simple — "use red/green TDD" is one of the most compact instructions you can give a coding agent that dramatically improves output quality.
The pattern is classic. Write tests first. Confirm they fail (red). Iterate implementation until they pass (green). Developers have been doing this for decades.
It's particularly effective for agents because it constrains two common failure modes: writing code that doesn't work, and writing code nobody asked for. The test gives the agent a concrete target. Requiring the test to fail first verifies it actually exercises the new code rather than passing by default.
Willison emphasizes something easy to overlook: confirming the red step matters. Skip it and you risk a test that already passes, which means your implementation was never validated. The agent will happily move on, and you'll have a false sense of coverage.
The Friction Is Real
Willison's guide doesn't fully address the practical difficulty of getting agents to follow TDD discipline.
Marc Love wrote about this directly. Even when you explicitly demand TDD in your CLAUDE.md or AGENTS.md, agents often ignore it. When they do attempt it, they tend to write the entire test file and the entire implementation at once. That's writing tests. It's not test-driven development. The red-green-refactor cycle requires incrementality, and agents default to doing everything in a single pass.
One developer built a multi-agent system specifically to solve this problem. Claude Code defaults to implementation-first. When you force TDD in a single context window, the implementation "bleeds" into the test logic — what he calls context pollution. His solution uses subagents where each agent starts with only the context it needs. He describes it as the only way he's found to get genuine test-first development from an LLM.
The Bigger Point
The deeper insight is about specification, not testing methodology.
Tests are specifications expressed in code. When you write a test first, you're defining what the system should do before anyone (or anything) builds it. The agent generates implementation to satisfy the spec. This is the same pattern at every scale — tests at the function level, architecture docs at the system level, domain models at the organizational level. Each one is a control surface for AI-driven engineering.
The teams getting the most from coding agents are the ones specifying well. Red/Green TDD is the entry point to that discipline.
Practical Takeaways
If you're working with coding agents today:
Put "use red/green TDD" in your CLAUDE.md or system instructions. It's compact and every major model understands it.
Watch for the cheat. Agents will write tests and implementation simultaneously if you let them. Enforce the red step.
For complex features, consider separating test-writing and implementation into distinct agent sessions to avoid context pollution.
Treat your test suite as a specification artifact. It's the clearest input you can give an agent about what "done" looks like.
Willison is building out a full guide series on agentic engineering patterns. This is worth following — he's assembling the practical playbook for how software actually gets built with agents.


