Todos los artículos

Part 2: The Transactional Outbox

Surviving a Broker Outage

15 de enero de 2026

#distributed systems#PostgreSQL#software architecture#Reliability#data-engineering
Part 2: The Transactional Outbox

In Part 1, we introduced the Observable Sandbox—a toy distributed system running in Docker that lets us safely simulate production failures. We used it to visualize backpressure by pausing a service.

Today, we are going to look at a more dangerous scenario: What happens when your infrastructure actually fails?

The Architecture Recap

As a reminder, here is how our data flows in the ingestion path:

We are focusing specifically on the relationship between the Ingestion Service, the Database, and RabbitMQ.

The "Dual-Write" Problem

A common mistake in distributed systems is the "Dual-Write." This happens when code tries to do two things:

  1. Save data to a database.

  2. Publish a message to a queue.

If you write db.save() then rabbit.publish(), and your app crashes in between, you have "ghost data" in the DB that no one knows about. If you reverse the order, you might send a message for data that never got saved.

The Illusion of Continuity and the Nefarious Scheduler

When we write code, we operate under a comforting illusion of continuity. We look at two consecutive lines in our IDE—line 40 saving to a database, and line 41 publishing an event—and we implicitly trust that if the first finishes, the second is inevitable.

But in the messy reality of production environments, running on networked infrastructure managed by container orchestrators, that trust is misplaced.

Between any two instructions lies an infinitely small, yet terribly dangerous, gap. This is the domain of the "nefarious scheduler." Whether it's the operating system kernel managing CPU slices, a Kubernetes node deciding your pod has consumed too much memory, or a sudden power fluctuation in a data center rack, outside forces can pause, preempt, or kill your process exactly at the worst possible moment.

This reality makes reasoning about distributed systems incredibly difficult, particularly when faced with the common "Dual-Write" requirement.

A dual-write occurs whenever a single logical operation needs to update two disparate systems that do not share a transactional boundary—for example, persisting business data to PostgreSQL and simultaneously announcing that change via RabbitMQ.

If you save to the database first, and the nefarious scheduler terminates your application in the microsecond before the message publishes, you have created "ghost data." Your database holds a truth that the rest of your system will never know about. If you reverse the order, publishing first and crashing before the save, downstream services begin reacting to data that doesn't actually exist in your system of record.

We cannot architect reliable systems based on the hope that the scheduler will be kind to us. We need a mechanism that acknowledges this chaotic reality. We require a way to bind the storage of facts and the communication of events into a single, atomic unit that either succeeds completely or fails cleanly, leaving no trace of partial execution.

The Solution: The Transactional Outbox

To fix this, our sandbox implements the Transactional Outbox Pattern.

  1. Ingestion Service: Receives data and writes it to a PostgreSQL table called outbox within a single SQL transaction. It does not talk to RabbitMQ directly.

  2. Dispatcher Service: A separate process that polls the outbox table. It picks up pending rows and publishes them to RabbitMQ.

  3. Completion: Only after RabbitMQ confirms receipt does the Dispatcher mark the row as DISPATCHED.

The Experiment: Kill the Broker

Let's prove this works by simulating a catastrophic failure of RabbitMQ.

Step 1: Verify Steady State

Make sure your stack is running (docker compose up -d). In Grafana, check the Outbox Status panel. You should see a steady flow of messages, and the "Pending" count should be near zero.

Step 2: The Outage

We are going to stop RabbitMQ. In a normal application, this would result in errors and lost data.

Run this command (or click Stop in the Docker Dashboard):



docker compose stop rabbitmq

Step 3: Watch the Build-up

Look at the Outbox Status panel in Grafana.

You will see the Pending count start to climb steadily.

I just stopped rabbit and in a few seconds, we have 1,080 message backed up.

Here is what is happening:

  • The Ingestion Service is still working perfectly. It’s writing data to Postgres. It doesn't care that RabbitMQ is down.

  • The Dispatcher is trying to send messages, failing to connect, and backing off. It leaves the records as PENDING.

Crucially, we are not losing data. The requests are being accepted and persisted safely to disk in Postgres.

Step 4: The Recovery

Let's bring the infrastructure back online.



docker compose start rabbitmq

Watch the dashboard. The Dispatcher will reconnect, detect the backlog, and flush the messages to RabbitMQ. The "Pending" line will drop, and the "Dispatched" count will spike as the system catches up.

Outside Reading

Focus: The critical center intersection. The diagram shows data being persisted to Postgres, and a separate "Dispatcher" polling that same DB to do further work. This is the visual representation of solving the "Dual Write Problem."

Diagram Scope: Ingestion Service → Postgres ← Dispatcher (Polling)

Topics to Investigate:

  • The "Dual Write" Problem: Why you cannot reliably save to a database AND publish to a message queue in the same service code block without risking data inconsistencies.

  • Distributed Transactions vs. Local Transactions: The difficulty of coordinating state across different infrastructure pieces.

  • The Transactional Outbox Pattern: The star of this architectural show.

    • What is an "Outbox" table?

    • The importance of atomicity: Committing the sensor data and the intent to process it further in the same DB transaction.

  • Polling Strategies (The Dispatcher):

    • How the Dispatcher finds unprocessed records in the Outbox table.

    • Naive polling vs. smarter strategies to avoid hammering the database.

    • Stretch topic: Change Data Capture (CDC) as an alternative to polling.

Summary

By using the Outbox pattern, we decoupled our data integrity from our infrastructure availability. We can turn off a critical component (the broker), and the system doesn't crash—it just pauses the distribution of data until things are healthy again.

In the final part of this series, we’ll look at how to debug this flow using Distributed Tracing and inject some network chaos using Toxiproxy.