Todos los artículos

The Agent Ops Stack, Explained: Two New Tools and the Layer They Belong To

2 de junio de 2026

#AI#Software Engineering#Devops##llmops#agents
The Agent Ops Stack, Explained: Two New Tools and the Layer They Belong To

Most discussions of AI infrastructure stop at the model. The model is the part that is easy to talk about — it has a name, a version number, and a benchmark score. The harder and less-discussed part is everything that sits around the model when it runs in production. That surrounding layer is starting to acquire a name. People are calling it the agent ops stack, and two projects released this week are good entry points for understanding what it is.

The first project is Golem.cloud, an open-source durable execution runtime aimed at agentic workflows. To see why something like Golem matters, it helps to think about what an agent actually does in production. A non-trivial agent rarely answers a single question and exits. It calls a tool, waits for a result, calls another tool, maybe asks a human for input, then resumes and writes some output. Each of those steps takes time. Some of them cost real money in tokens or API fees. And every step is a place where the process can crash, the server can be redeployed, or the network can drop a connection.

This problem is not new. Traditional distributed systems solved it in the 2010s with the pattern called durable execution, popularized by systems like Temporal and Cadence. The pattern records the state of a workflow at each step so that a failed run can resume from where it left off rather than starting over. Golem takes that same idea and retargets it specifically at agent workflows, where the cost of restarting a 30-step research run is much higher than the cost of restarting a typical microservice call.

If you have been building agent systems on top of a Postgres table and a queue of jobs, Golem is worth understanding as a category. It does not necessarily replace what you have built, but it gives you a vocabulary for what category your custom solution belongs to and a reference point for what features it might be missing — things like deterministic replay, versioning of workflow definitions, and built-in handling of human-in-the-loop pauses.

The second project is Tokentoll, a continuous integration tool that tracks LLM API costs on a per-pull-request basis. The premise is simple: a single prompt template change or a tweak to an agent's loop can multiply token usage by a factor that does not show up in any test until the next billing cycle. Tokentoll runs your test suite, records the tokens consumed, compares the result to a baseline, and fails the build if the cost regression exceeds a threshold you set.

The analogy here is to performance budgets in front-end engineering. The bundle-size check that fails when a new dependency pushes a JavaScript bundle over a size limit is a familiar pattern. The same idea applied to inference cost is overdue. As more application logic moves into prompts and agent loops, more of your production cost moves into a line item that traditional CI never measured. Putting a gate on that line item closes a gap that has been quietly widening for two years.

Looking at these two tools together is more instructive than looking at either one alone. They occupy the same conceptual layer: not the model, not the framework that calls the model, but the operational substrate that determines whether your agents survive contact with production and whether the bill stays inside the budget you committed to. That layer is starting to look a lot like the operating layer that DevOps built for traditional services in the 2010s — durable runtimes, cost observability, deployment patterns, rollback strategies, replay tooling.

It is reasonable to expect that the next year or two will produce more entries in this category, and that the patterns will start to standardize. If you have not been thinking about your AI stack in these terms, the practical exercise worth doing this week is to draw your current architecture on a single page and label each layer: model, framework, agent harness, runtime, cost control, evaluation. Note which layers you have invested in and which you have left to ad-hoc scripts. The gaps that show up on that page are usually the ones that produce the next quarter's outages.

Tools like Golem and Tokentoll are useful in their own right, but the more durable takeaway is the shape of the stack they imply. That shape is what you will be building on top of for the next several years.