All posts

The Most Important File in Your Repository is domain.md

February 17, 2026

#AI#llm#agentic AI#Developer
The Most Important File in Your Repository is domain.md

Image Credit: Oleg Gapeenko

I’ve spent roughly twenty years leading engineering teams—as a CTO, VP of Engineering, Head of Product Engineering, etc. Now I run my own firm helping companies accelerate timelines and reduce risk in delivery and help turn around teams. One benefit of running a firm is that I get dropped into new business domains all the time. It's great. I love it and learn a lot. Here's a recent lesson I learned that improved the quality of code my agents are creating and am happy to share.

Early in my career, I over-indexed on pure technical talent. If someone could reason about distributed systems, refactor code cleanly, or implement an elegant abstraction, I was impressed. And to be clear, technical excellence matters. But over time, something became obvious: the engineers who consistently made the best decisions weren’t just technically strong. They understood the domain.

They understood:

  • The domain-specific terminology and concepts AND could communicate seamlessly with other domain SMEs.

  • How the company made money.

  • Why customers bought.

  • Where risk lived.

  • How operations actually worked.

  • What failure meant.

That context has a material impact. Engineers who grasp the business domain have a much higher chance of building the right product the first time. They earn the respect of subject matter experts because they speak the same language, which in turn builds trust between the business and its engineering team. When an engineer understands the business deeply, their local decisions align with global outcomes. Absent that context, even strong engineers drift toward generic solutions.

Why Context Improves Decision Quality

From first principles, this is straightforward. Decision quality depends on the information available. Under the theory of bounded rationality, agents (human or otherwise) operate with limited knowledge and make the best choice they can from what they see. Experts outperform novices not simply because they are smarter, but because they possess domain-specific schemas that help them recognize patterns and filter irrelevant noise. In team settings, shared mental models correlate with higher performance because alignment reduces coordination errors.

In short:

Context improves reasoning. Remove context, and you increase variance.

Now, consider AI agents. Large Language Models (LLMs) are probabilistic pattern engines. They don’t “understand” your business unless you encode that understanding for them. If you don’t provide domain constraints, they default to internet-average SaaS assumptions. Which is fine—until you’re not building average SaaS.

AI Makes Context Management Non-Optional

As I began working deeply with AI-based coding systems, it became very clear: context management is the core discipline of AI-driven engineering.

Greenfield projects often move fast because the domain context is small and the core assumptions can fit inside the model’s context window. The AI can hold most of the system in its metaphorical memory. But as systems grow, that context gets compressed. Agents summarize, they forget, they collapse nuance, and they revert to defaults. Once that happens, you start seeing:

  • Generic architectural patterns

  • Soft security assumptions

  • Casual data modeling

  • UI tone that doesn’t match the risk posture

  • “Happy path” implementations in domains that require rigor

This is a solvable context problem, and the solution is becoming a non-negotiable asset for modern software teams.

The Research is Clear: Grounded LLMs Perform Better

Multiple studies and industry best practices confirm that grounding AI tools in domain-specific context dramatically improves their performance. A recent research paper found that general-purpose LLMs struggle with domain-specific code tasks because they lack familiarity with specialized libraries and APIs. In the study, simply prompting the model with relevant API documentation led to “more professional code” and better solutions.

Industry anecdotal evidence echoes this. Google engineer Addy Osmani stresses that “LLMs are only as good as the context you provide – show them the relevant code, docs, and constraints”. This is because a well-informed model can understand project-specific constraints—performance requirements, error handling conventions, etc.—and make choices aligned with your domain. Engineers even have a term for this: “context engineering,” which focuses on “bringing the right information (in the right format) to the LLM” to get more accurate and reliable code.

This principle is already standard practice at top engineering organizations, which have long recognized that deep domain understanding is critical for human developers. Stripe is famous for its comprehensive internal docs that get new engineers up to speed quickly. Google now places new engineers on teams with an AI chatbot trained on Google’s vast corpus of internal documentation, effectively serving as an always-available mentor.

These practices underscore a foundational principle: documented context is not just for humans—it’s fuel for AI reasoning. When we give AI agents these artifacts, we are teaching them the same way we teach new team members.

domain.md: Leadership Formalized

For years, I saw it as part of my job to teach engineers the domain. I would study the industry, talk to experts, understand revenue mechanics, and map operational constraints. Then I would teach that to my team. If you've worked with me in recent years, you'll have seen this pattern first-hand. Today, I encode that discipline into a file: domain.md.

This file is contextual grounding for decision-making, the background layer for both human engineers and AI agents. It contains:

  • The precise domain the business operates in

  • How the company makes money and why it exists

  • The “product core” (the 5-10 conceptual entities everything maps back to)

  • Operational constraints and risk posture

  • What the company explicitly is not

  • Market context and competitors

An Example domain.md File



# domain.md ## PortLedger This document exists to ground engineers and AI coding agents in the exact business domain PortLedger operates in. It is not marketing material and not a technical specification. It defines the domain context that should shape all architectural and product decisions. ## Domain PortLedger operates in the commercial maritime logistics domain. Specifically, it provides operational coordination software for mid-to-large cargo ports that manage containerized freight, bulk shipments, and intermodal transfers. This is a real-world, asset-heavy, time-sensitive environment. Delays cost money. Mistakes can halt vessel operations and create contractual penalties. The domain involves: * Port authorities * Terminal operators * Shipping lines * Trucking companies * Customs and regulatory agencies * Yard operations staff ## Core Problem Ports coordinate thousands of physical movements daily: * Vessel arrivals * Container unloading * Yard allocation * Customs clearance * Truck gate scheduling * Intermodal rail transfers Most mid-tier ports still rely on fragmented systems, spreadsheets, email chains, and phone coordination. The structural problem is not “data entry.” The structural problem is real-time coordination of physical assets across independent organizations with financial and regulatory consequences. Failure results in: * Demurrage fees * Missed vessel windows * Yard congestion * Regulatory violations * Revenue leakage PortLedger exists to make operational coordination explicit, auditable, and financially aligned. ## How PortLedger Makes Money PortLedger operates as a B2B SaaS platform. Revenue model: * Annual contracts with port authorities or terminal operators * Tiered pricing based on container volume * Optional paid modules (advanced analytics, predictive congestion modeling) Retention depends on: * Reduced vessel turnaround time * Improved yard utilization * Fewer scheduling conflicts * Measurable cost savings If the platform increases throughput or reduces penalties, customers renew. ## What PortLedger Does PortLedger provides a centralized coordination platform that: * Tracks vessel schedules and berth assignments * Manages container inventory within the yard * Coordinates truck gate appointments * Tracks customs clearance status * Provides operational dashboards for supervisors * Logs audit events for operational decisions It connects operational actors who do not share a single system of record. The system must represent physical reality accurately and in near real time. ## Operational and Regulatory Constraints * Customs compliance requirements vary by country. * Cargo manifests contain regulated data. * Operational decisions may have contractual consequences. * Data accuracy is financially material. * Downtime during peak vessel arrival windows is unacceptable. * Time zones and cross-border communication are normal. The system cannot assume perfect connectivity. It must tolerate partial failure and asynchronous updates. ## The Product Core All architectural reasoning should anchor to these conceptual entities: * Vessel * Berth * Container * Yard Slot * Movement Event * Gate Appointment * Organization * Contract * Customs Status * Audit Event These are conceptual nouns, not database tables. If a new feature cannot be mapped back to these core concepts, it likely does not belong in this system. ## What PortLedger Is Not PortLedger is not: * A generic project management tool * A document storage system * A billing platform * A GPS tracking service * A predictive AI-only analytics company It is an operational coordination platform anchored in physical asset movement. ## Risk Posture Primary risks: * Incorrect container location data * Scheduling conflicts causing vessel delays * Unauthorized access to cargo manifests * Failure to record operational audit events * Data loss during peak load Failure is measured in operational disruption and financial penalties. The system must favor correctness and auditability over UI novelty. ## Design Implications Architectural and product decisions should assume: * Multi-organization access control * Strict event logging * Clear separation of operational states * Idempotent movement updates * Explicit timestamps with timezone awareness * Conservative UI patterns in operational dashboards “Move fast and break things” is incompatible with this domain. This is the finished `domain.md`. It is: * Precise * Operationally grounded * Business-aware * Blunt * Free of marketing fluff Now here’s the key question. Do you want to: 1. Run the “same feature, different domain” experiment using this PortLedger example vs. a lightweight consumer SaaS example? 2. Or generate a second contrasting domain.md right now to pair with it? If we do that, the delta will be extremely clear.

How to Create a domain.md File

Here is the exact, six-step process I use when inheriting or starting a project.

Step 1: Open a Research Chat

Start with a clean chat thread with an AI assistant. Be explicit about your goal:

"I am creating a file called domain.md. Its purpose is to ground my AI coding agents and engineering team in the exact business domain this company operates in. I'm new to this domain, so you will be teaching us."

Step 2: Force Research

Instruct the model to analyze the company’s website, research competitors, identify the market structure, explain revenue mechanics, and identify regulatory or operational pressures. If the company has no public presence, you must provide this context manually. The goal is precision.

Step 3: Lock in the Boundaries

Force the model to define what the company does and, just as importantly, what it does not do. What problems does it explicitly not solve? What risks exist? What does failure look like? Ambiguity is the enemy of good AI outputs.

Step 4: Define the Product Core

List the 5–10 conceptual nouns that anchor the entire system (e.g., User, Order, Policy, Audit Event). These are not database tables; they are the core concepts that anchor architectural reasoning.

Step 5: Edit Ruthlessly

Remove all marketing fluff, empty adjectives, and vague language. Make it blunt, specific, and useful.

Step 6: Store and Reference

Commit domain.md to the repository. Reference it in agent system prompts. Share it with your engineers. Treat it as a living artifact. It is now part of your context architecture.

The Shift to Context Architecture

Senior engineers with domain awareness make better decisions. AI agents with domain awareness make better decisions. The variable is context.

Grounding AI assistants yields higher-quality, more relevant, and safer results. It transforms the AI from a clever guesser into a knowledgeable collaborator whose outputs align with your codebase, business logic, and safety requirements. In AI-native engineering, leadership increasingly looks like context architecture.

The most important file in your repository may not be main.py.

It may be domain.md.