記事一覧

Part 3: Governing the Inner Loop

Fast Probabilistic Guardrails for Autonomous Agents

2026年9月22日

#AI#Software Engineering#agents#architecture#Devops
Part 3: Governing the Inner Loop

Why Agent State Machines Belong in Code, Not Conversational Self-Reflection

By Jason Vertrees
Part 3 of the Machine-Native AI Series

The Autonomous Tool-Calling Trap

When building autonomous coding agents or enterprise workflow runners, the most dangerous architectural surface is the tool execution loop.

An agent operating with filesystem, terminal, or database access will execute dozens of commands during a single task. Inevitably, situations arise where a command could cause catastrophic, irreversible damage:

  • Running a hard git reset that wipes uncommitted working files.

  • Executing a database truncation or table drop.

  • Deleting files or altering configs outside the active task's declared scope.

To prevent this, the prevailing industry trend has been "LLM self-reflection": adding instructions to the agent's system prompt telling it to think about whether its pending tool call is safe before executing.

This is a fundamental design flaw. An agent that is hallucinating a dangerous action will happily hallucinate that the action is completely safe. English prose in a system prompt cannot enforce an invariant boundary.

Yet replacing self-reflection with static regex denylists is equally brittle. A pure regex cannot tell whether rm -rf tmp/ is an authorized cleanup within the task's scope or an accidental deletion of production source. Worse, when constrained, agents routinely smuggle shell execution through inline scripts (python -c "import shutil; ..."), which trips up or evades keyword matchers completely.

To govern an agent safely without grinding its iteration speed to a halt, you need rapid, sound judgment calls returning calibrated probabilities that deterministic code can evaluate.

The Architecture: The Calibrated Agent Dispatcher

Instead of letting the agent govern itself, we place a deterministic finite state machine between the agent's pending instruction and the execution environment:

               Incoming Agent / User Instruction
                               │
                               ▼
               ┌───────────────────────────────┐
               │   TypeSafe System One (Jev)   │
               │   • is_destructive   (Noul)   │
               │   • next_transition  (Choice) │
               │   • clarity_score    (Score)  │
               └───────────────┬───────────────┘
                               │
                   Calibrated Probabilities
                               │
                               ▼
               ┌───────────────────────────────┐
               │    Deterministic State Machine │
               │    (Finite Python Transition)  │
               └───────────────┬───────────────┘
                               │
      ┌────────────────────────┼────────────────────────┐
      ▼                        ▼                        ▼
[ AUTO_EXECUTE_READ ]  [ TRIGGER_2FA_CONFIRM ]  [ REJECT_AMBIGUOUS ]
  p_destr < 0.10          p_destr >= 0.95         clarity < 3.0
  clarity >= 3.0          Requires Human          Halts Runaway Loops

Here, the model acts strictly as a calibrated sensor. It measures:

  1. is_destructive: Calibrated probability that the action causes irreversible state loss.

  2. next_transition: Categorical classification (READ_DATA, UPDATE_RECORD, EXECUTE_DESTRUCTIVE_ACTION, REQUEST_CLARIFICATION).

  3. clarity_score: Ordinal measurement (1 to 5) of how complete and unambiguous the instruction is.

The State Machine Implementation

Your application code enforces strict, testable mathematical transitions:

from typesafe_sdk import Choice, Noul, Score

DISPATCHER_QUESTIONS = {
    "is_destructive": Noul(
        instructions="Does the pending instruction request a permanent deletion, destruction, or irreversible mutation?"
    ),
    "next_transition": Choice(
        instructions="Determine the action category of the pending instruction.",
        criteria={
            "READ_DATA": "Querying, reading, or inspecting data without modifying state",
            "UPDATE_RECORD": "Modifying, inserting, or configuring standard non-destructive records",
            "EXECUTE_DESTRUCTIVE_ACTION": "Permanent deletion, purging, dropping tables, or destructive resets",
            "REQUEST_CLARIFICATION": "Vague, ambiguous, or incomplete instructions requiring operator guidance",
        },
    ),
    "clarity_score": Score(
        instructions="Rate the clarity and completeness of the instruction from 1 to 5.",
        criteria=[
            "Level 1: Completely ambiguous or missing context (e.g., 'fix the thing')",
            "Level 2: Vague intent, missing critical target parameters",
            "Level 3: Moderate clarity, primary intent identifiable",
            "Level 4: Clear intent with well-specified target resources",
            "Level 5: Crystal clear, explicit parameters with no ambiguity",
        ],
    ),
}

The router logic translates these calibrated outputs into deterministic actions:

def route_agent_action(state: dict) -> DispatcherDecision:
    with client:
        res = client.system_one(state=state, questions=DISPATCHER_QUESTIONS)

    p_destr = res.nouls["is_destructive"].noul
    transition = res.choices["next_transition"].choice
    clarity = 1.0 + res.scores["clarity_score"].score  # Normalize 0-4 to 1-5

    # 1. High-confidence destructive action -> Enforce human approval
    if p_destr >= 0.85 or transition == "EXECUTE_DESTRUCTIVE_ACTION":
        if p_destr >= 0.95:
            return DispatcherDecision(
                action="TRIGGER_2FA_CONFIRMATION",
                reason="High-confidence destructive intent. Human confirmation required.",
            )
        # Destructive action with uncertain parameters -> Abort to prevent disaster
        return DispatcherDecision(
            action="HARD_ABORT_UNCERTAIN_DESTRUCTION",
            reason="Destructive intent detected but parameters are uncertain.",
        )

    # 2. Ambiguous or underspecified prompt -> Halt and prompt user
    if clarity < 3.0 or transition == "REQUEST_CLARIFICATION":
        return DispatcherDecision(
            action="REJECT_AND_REQUEST_CLARIFICATION",
            reason=f"Instruction clarity below threshold ({clarity:.1f}/5.0 < 3.0).",
        )

    # 3. Low-risk read or query -> Dispatch automatically without human interruption
    if transition == "READ_DATA" and clarity >= 3.0:
        return DispatcherDecision(
            action="AUTO_EXECUTE_READ",
            reason="Safe read operation with high clarity. Dispatching automatically.",
        )

    # 4. Standard in-scope mutation
    return DispatcherDecision(action="EXECUTE_STANDARD_MUTATION")

Empirical Behavior in the Agent Loop

When tested against realistic agent command scenarios, the dispatcher separates routine exploration from destructive accidents without false positives:

Incoming Agent Command Calibrated Sensor Reading Deterministic State Transition
"Fetch total active enterprise users in EU-West for Q3" p_destr = 0.01
trans = READ_DATA
clarity = 4.7 / 5.0 AUTO_EXECUTE_READ
(Dispatches immediately)
"Purge all inactive customer records from production db immediately" p_destr = 0.98
trans = EXECUTE_DESTRUCTIVE
clarity = 4.8 / 5.0 TRIGGER_2FA_CONFIRMATION
(Blocks execution pending human approval)
"Fix the broken report from yesterday" p_destr = 0.04
trans = REQUEST_CLARIFICATION
clarity = 1.9 / 5.0 REJECT_AND_REQUEST_CLARIFICATION
(Stops agent before it guesses wrong)
"Run database cleanup on test accounts" p_destr = 0.88
trans = EXECUTE_DESTRUCTIVE
clarity = 2.4 / 5.0 HARD_ABORT_UNCERTAIN_DESTRUCTION
(Catches ambiguous high-blast-radius command)

Architectural Takeaway

The holy grail of agent architecture is maximum autonomy within provable safety boundaries.

By replacing conversational self-reflection with rapid, calibrated judgment calls, your harness lets agents move at full speed on routine reads and safe edits, while catching destructive operations and ambiguous instructions before damage occurs.

In Part 4, we scale this pattern to DevOps pipelines: using calibrated failure probabilities to distinguish test flakes from real code regressions in real time.


Jason Vertrees is the founder of Heavy Chain Engineering, which helps lower middle-market vertical SaaS companies and PE firms turn scattered AI usage into measurable delivery leverage — 85% faster feature velocity, six-to-eight-week projects shipped in days. If you want help building an AI-native engineering organization, book an AI Delivery Assessment or email jason.vertrees@gmail.com.