
TL;DR
A backend test suite started "hanging" at 99% completion. Two days later, with help from py-spy, pytest-timeout, and a JUnit XML report, we found four root causes — three real test failures and one SQLAlchemy LRU cache that thrashed under Python 3.14. None of them were a hang. The agent (an LLM coding assistant) and I were stuck in a tight loop of "try another pytest command, kill, repeat" for the better part of forty hours. The technical post-mortem below covers the bugs. The harder lesson — and the one I want to leave you with — is about how to recognize that you and your tools are in a degenerate debugging loop before you've burned two days finding out.
Setup
Mid-sized B2B platform. Python 3.14.2 backend, FastAPI, async SQLAlchemy on Postgres, around 12,000 backend tests. We use pytest-xdist with --dist loadfile for parallelism, a pre-push hook that runs lint + the full suite + a small E2E layer, and a 98%-coverage gate. On a healthy day the suite completes in 2–3 minutes on a developer laptop with 18 workers.
This is a healthy harness — the kind that catches things. Most of the time, it does.
The Surface Failure
I tried to push a small batch of bug fixes. The pre-push hook ran pytest, and after a few minutes the dots stopped advancing. The run sat at 99% for what felt like a long time. I killed it and tried again. Same thing. I tried -x (stop on first failure). The first failure landed at 84%, then the suite hung again — different behavior, same outcome.
That was the opening move of a marathon.
The Wrong Paths
The most expensive thing about this debugging session was not any single mistake. It was cycling through plausible diagnoses without ever stopping to ask whether the next attempt was materially different from the last. A partial list:
"It's a slow test." Increase the timeout. The suite still hung.
"It's xdist's per-worker DB provisioning racing." Add an advisory-lock guard, restart, hang.
"The session-scope autouse cleanup fixture is killing live sibling-worker connections." True, in retrospect, but fixing it didn't unstick the run.
"It's coverage instrumentation slowing things down past my patience." Pass
--no-cov. Same hang."Let me see what one test is doing" — pass
-xvsto the actual hanging file. Pytest dumped thousands of lines of audit log noise, the terminal scrolled, I assumed it was hung, and I killed it. (It wasn't hung. The-sflag had unmasked aWARNING-level audit log that fires every time a test fixture inserts a row without anorganization_id— exactly the strict-tenant isolation rule firing correctly.)"Use
--lfto rerun only what failed last time." The.pytest_cache/v/cache/lastfailedfile was poisoned by every previous interrupted run.--lfreported 98 failures, started replaying them, and hung again at the first slow file."OOM-killed?" I had the agent check
log showforJetsamevents. 128 GB of RAM, no Jetsam events, no Python crash reports in~/Library/Logs/DiagnosticReports. Workers were not being killed."Workers vanished without crashing." A bad reading of
psoutput —pytest-xdistworkers spawn aspython -u -c <bootstrap>, not as anything matchinggrep pytest. They were there the whole time; I'd been grepping for the wrong string.
By the end of day one I had run pytest dozens of ways, killed the controller eight or nine times, and had a strong feeling something was genuinely broken — but no specific diagnosis.
The Diagnostic Chain That Actually Worked
The breakthrough was a sequence of three tools, used in order.
1. py-spy dump on the still-running controller
py-spy is a sample-based profiler that doesn't need to be linked into the target process. On macOS it needs sudo. The first dump showed the pytest controller's main thread parked in xdist/dsession.py:154, waiting on its event queue, and 18 receiver threads parked in execnet/gateway_base.py:534 reading from worker pipes. Some of the workers were idle (sitting on read() waiting for new work); one worker had its main thread active+gil and the top of its stack was inside SQLAlchemy's compiled-statement LRU cache:
_inc_counter (sqlalchemy/util/_collections.py:513)
get (sqlalchemy/util/_collections.py:528)
_compile_w_cache (sqlalchemy/sql/elements.py:712)
_execute_clauseelement (sqlalchemy/engine/base.py:1633)
...
execute (sqlalchemy/orm/session.py:2351)
That was the first hard datum in two days. The "hang" was not a hang at all. It was one worker, running real Python, churning through SQLAlchemy's LRU cache eviction. Default query_cache_size=500, a 12,000-test suite that emits a wide variety of compiled statements, and Python 3.14's adaptive interpreter all combined to make LRUCache._inc_counter (the per-access counter bump) visible in the profile.
2. py-spy record for a flamegraph
A two-second dump tells you what the process is doing right now. A 30-second flamegraph tells you what it's been doing in aggregate:
py-spy record --pid <worker_pid> --duration 30 --output flame.svg
The flamegraph put 9% of all samples in a single test function — test_backfill_processes_vendor_profile_physical, with another 9% spent inside the production function it was calling. (Names sanitized here.) That's the test. That's the file the loadfile scheduler had pinned to that worker. The rest of the workers were idle waiting for the slow one to finish so the controller could dispatch more files.
3. pytest-timeout + --junit-xml
Once I knew the suite wasn't actually hung — just slow on one or two tests — the right move was to put a time limit on individual tests and let the slow ones self-kill so the suite could report the rest:
uv pip install pytest-timeout
pytest -n auto --dist loadfile \
--timeout=90 --timeout-method=signal \
--junit-xml=/tmp/junit.xml \
--tb=short --no-cov -q
--timeout-method=signal uses SIGALRM, which interrupts the main thread even if it's blocked in C extension code. --junit-xml writes test results incrementally — even if you have to Ctrl-C the run, you can parse what's already on disk and see every test that completed.
Three minutes later, real results:
11,873 passed
4 failed: two timeouts on backfill tests, one stale assertion, one schema-drift check tripping on test-fixture ORM models that had leaked into the global
Base.metadata.21 skipped, 1 xpassed.
What Was Actually Wrong
Four distinct issues, none of them a hang.
1. SQLAlchemy compile-cache thrash. SQLAlchemy's default query_cache_size=500 is sized for typical application workloads. A test suite that exercises many code paths emits many more distinct compiled statements than that. The LRU thrashes; per-test cost rises non-linearly with suite size. Fix: one line in the test engine factory — query_cache_size=5000.
2. An infinite loop in a paginated backfill drain. This is the prize. A recent feature changed a batch task from a single-page select to a "drain until empty" while True loop:
while True:
ids = await _select_ids_needing_backfill(...)
if not ids:
break
for record in load(ids):
input = _extract_input_from_jsonb(record.payload)
if input is None:
continue # ← row's FK is never updated
...
await _set_fk(record.id, ...)
await session.flush()
Now consider what happens when the JSONB column was written from a Python None. SQLAlchemy doesn't always map Python None to SQL NULL on JSONB columns — under common configurations it stores 'null'::jsonb. The Postgres predicate column IS NOT NULL is true for a JSONB null, so the same row keeps matching the select forever. _extract_input_from_jsonb(None) returns None, the loop hits continue, the row's FK never gets set, the next iteration's select returns the same ID set. Infinite loop in test (timed out at 90s); infinite loop in production (a Celery worker that would have spun on the first row it couldn't extract input from).
The fix is small but the failure mode is brutal:
while True:
ids = await _select_ids_needing_backfill(...)
if not ids:
break
before = counters.normalized
...for record in ...:
...
await session.flush()
if counters.normalized == before:
# No row in this page was processable; the select would
# return the same IDs next pass. Break instead of looping.
break
This is a real production safety improvement, not just a test fix. A single row with an unprocessable JSONB payload would have wedged the Celery beat job.
3. A stale test assertion. A recent commit added a category field to an error-record dict; the test that asserted shape equality hadn't been updated. Standard staleness, trivial fix.
4. Test-fixture ORM models polluting the global registry. Three test files defined ORM models with __tablename__ starting with test_ for fixture use (a Widget, a RepoItem, an EntityModel). At module import time these get registered on the shared Base.metadata, so a later "ORM-vs-Alembic schema drift" CI check iterated them and tried to SELECT against tables that exist only in test imagination. The drift check now filters out tables whose name starts with test_. The cleaner long-term answer would be to give test fixtures their own DeclarativeBase, but the filter is a one-line fix and is correct in spirit: the global registry was polluted by test code's import side effects.
Tools That Earned Their Keep
For anyone who wants to replicate the diagnostic path:
py-spy(py-spy dump,py-spy record) — sample-based, no instrumentation, attaches to a running process. Needs root on macOS because oftask_for_pid.dumpfor "what is this process doing right now";record+ flamegraph for "what has this process been doing for the last 30 seconds." On Python 3.14 the unwinder is occasionally flaky; if a dump returns empty output, try again.pytest-timeoutwith--timeout-method=signal. Per-test wall-clock cap; on timeout, the suite continues rather than hanging. Thesignalmethod is critical for async tests —threaddoesn't kill coroutines waiting on async I/O.pytest --junit-xml=for survivable test output. Writes results incrementally, so killing the run preserves everything that finished. Parse it after the fact withxml.etree.ElementTreefor failure summaries.pytest --lf(last failed) with the caveat that the.pytest_cache/v/cache/lastfailedfile persists across runs. Wipe the cache (rm -rf .pytest_cache) before relying on--lfafter a series of interrupted runs.Reading
psoutput forpythonrather thanpytest— pytest-xdist workers don't havepytestin their argv; they're spawned aspython -u -c <bootstrap>.pkill -INTthen-9in sequence for cleanly killing pytest controllers and workers.-INTgives pytest a chance to flush results;-9cleans up if it doesn't.
Pytest-xdist Behavior Worth Internalizing
Two gotchas in pytest-xdist were responsible for most of my wrong-path debugging.
Post-failure drain. When you pass -x and a worker reports a failure, the controller stops dispatching new test files but lets in-flight workers finish their current file. Under --dist loadfile some files contain slow integration tests, so the drain can take many minutes. From the outside this looks identical to a hang. It isn't. If you Ctrl-C, you lose the failure traceback (it's only printed at the end of the run).
Workers exit silently. If a worker process is SIGKILLed (OOM, crash, manual kill -9), it does not send its session-complete sentinel through execnet. The controller waits forever for that sentinel. From the outside the controller looks hung. The ps output will show no python children of the controller — that's your tell.
Neither of these is a bug in xdist. Both are reasonable behaviors given its design. They just need to be in your head when you read the symptoms.
The Meta-Lesson
The technical lessons above are fine. Useful, even. The real lesson is structural.
The AI coding assistant and I were in a tight loop for two days. Every turn, the model generated a plausible-sounding next pytest command. Every turn, I ran it. Every turn, the command produced an ambiguous result — it took longer than I wanted, or it produced output I half-read, or it hung at exactly the same point as the previous attempt. The model interpreted each new result as a new clue and generated a new variant. I was so primed to expect progress that I kept reading the new variant as progress.
It wasn't. We were generating low-information moves at high frequency — the debugging equivalent of churn in a poorly-designed feedback control loop. The agent did not detect this. I did not detect this, not until I had crossed a personal threshold of frustration and asked the agent point-blank: "How did we let this happen?"
The honest answer was: I had been treating the agent's confident tone as evidence that the agent had a plan. The agent had no plan. It had patterns. It was matching the pattern "user reports a hung pytest" to the pattern "suggest a more diagnostic pytest invocation." It was doing this with high fluency. It was wrong every single time.
The cost: roughly two days of senior-engineer time, several thousand dollars in fully-loaded labor, blocked roadmap progress, and a personal frustration cost I'd prefer not to repeat.
The thing that actually broke the loop was nothing technical. It was me stopping, taking a breath, and explicitly asking the agent to dispatch a different model for adversarial review. That model ignored most of the prior context, looked at the diff fresh, identified the autouse-cleanup fixture as the root cause within fifteen minutes, and we moved on. Then py-spy came in, then pytest-timeout, and the actual bugs surfaced.
What I Would Build Now
If I were building agent tooling from the ground up after this session, I would add one component before anything else.
A stuck-loop detector. A small hook that runs at the end of every turn, tracks (turns since last commit, turns since last new file modified outside __pycache__, similarity between the current Bash command and the previous three Bash commands), and emits a structured alert when the ratio of productive-to-total turns collapses past a threshold. The alert needs to go in two directions:
Into the agent's next-turn context, with explicit instructions: re-baseline, kill in-flight processes, change approach, escalate to an adversarial-review model.
Onto the operator's screen, with a visible "we are not making progress" banner. This is the part the human needs. Without it, the operator's only signal is emotional impatience, and emotional impatience is a slow and unreliable trigger when the agent's confident-but-wrong output keeps suggesting "one more attempt."
"How is the agent doing?" is a question the harness should be able to answer, not a question the operator has to feel their way to.
Closing
The four bugs in the test suite were genuinely instructive — the JSON-null infinite loop in particular is a real production-safety issue that an agent-driven coverage push had failed to surface. I'm glad they came out. I am not glad about the time it cost to find them.
If your team uses AI coding assistants for debugging, the question to ask isn't "is the agent smart enough to find this bug." It's "would either of us notice if we were in a degenerate loop for two days?" Today, in most setups, the answer is no. Building the alarm for that is the most valuable single piece of harness tooling you can put in place this quarter.


