記事一覧

A Text File Beats an Intelligent Agent (And the Data Proves It)

2026年4月1日

#AI#Software Engineering#devtools#agentic AI#Open Source
A Text File Beats an Intelligent Agent (And the Data Proves It)

Image by Susan Wilkinson

Two competing approaches to giving AI coding agents up-to-date knowledge collided this week. Google released official "Agent Skills" for the Gemini API. Vercel published data showing a simple AGENTS.md text file achieves 100% on the same tasks — beating Google's more complex system.

The implications go deeper than a benchmarking anecdote.

Google's Approach: Agent Skills

Google shipped official skill packages for Gemini API development, Vertex AI, Gemini Live API, and the Interactions API. These are curated knowledge bundles the agent can invoke when it needs documentation. Gemini 3.1 Pro jumped from 28.2% to 96.6% on 117 coding tasks. Published on GitHub at google-gemini/gemini-skills.

96.6% is impressive. But the Vercel data makes it look like an unnecessary detour.

Vercel's Finding: Agents Ignore the Skill System

Vercel tested agents with access to Google's skill system and found something revealing: agents ignored the skills in 56% of test cases. They never called the skill system at all. Just went ahead and tried to solve the problem without looking anything up.

Pass rate with skills available but not forced: 53%. That's the same as having no documentation at all.

Only when researchers forced invocation — literally telling the agent "You MUST invoke the skill" — did the pass rate climb to 79%. The agent had to be coerced into using the tool designed to help it.

Then Vercel tried something simpler. They compressed the relevant documentation into an 8KB AGENTS.md file (down from 40KB of skill packages) and pre-loaded it into context.

100% success rate. Zero decision points. Zero sequencing issues. The knowledge was just there.

Why Passive Context Wins

The finding is counterintuitive if you think of agents as intelligent decision-makers. Shouldn't an agent that can choose to look things up outperform one that has everything dumped into context?

In practice, no. Every decision point is a failure point. The agent has to decide whether to look something up, what to look up, when in its workflow to look it up, and how to interpret what it finds. Each decision has a nonzero error rate. Stack enough of them and the cumulative failure probability overwhelms the benefit.

A static text file eliminates all of those decisions. The knowledge is pre-loaded. The agent reads it once and proceeds. No retrieval strategy. No tool-calling overhead. No sequencing errors.

This is the spec-driven development thesis in action: the specification shapes behavior deterministically rather than probabilistically.

AGENTS.md Has Already Won (For Now)

AGENTS.md is now adopted by over 60,000 open-source projects and natively supported by Cursor, Devin, GitHub Copilot, and Gemini CLI. It's the de facto standard for project-level agent context.

But there's a scaling problem.

lat.md: What Comes Next

A new project called lat.md launched on Hacker News this week with a direct thesis: AGENTS.md works for small projects but breaks down as codebases grow. The creator is Yury Selivanov — EdgeDB founder, Python core contributor.

The approach: a lat.md/ directory at your project root containing interconnected markdown files with [[wiki links]]. Source code links back via // @lat: comments, creating a bidirectional graph between documentation and implementation.

lat check validates all links resolve and all code references exist. The docs stay in sync automatically. Semantic search via embeddings, exact search via CLI, and an MCP server for editor integration.

The key insight: for complex codebases, flat context doesn't scale. You need structured context — a knowledge graph that both humans and agents can navigate. The evolution path looks like this: flat AGENTS.md for small projects, wiki-linked knowledge graphs for large ones, and eventually auto-generated specs that update as code changes.

What This Means For Your Team

If you're building with AI coding agents and haven't adopted AGENTS.md yet, start there. It's trivially easy and the data shows it works.

If you're running a large codebase and finding that a single context file isn't enough, watch lat.md. The pattern — structured, validated, bidirectional documentation — is where this is heading.

And if you're evaluating Google's Agent Skills or similar tool-based knowledge systems: the Vercel data suggests you should exhaust passive context approaches first. Give the agent the knowledge upfront. Only reach for active retrieval when the context budget can't hold everything the agent needs.

The simplest approach works best until it doesn't. Know when you've hit that boundary.


Jason Vertrees is founder and CTO of Heavy Chain Engineering, an AI-native software consultancy specializing in harness engineering, AI-driven SDLC, and fractional CTO services for teams scaling with AI.

Happy thinking, Jason