150 compound-AI incidents mapped to resilience patterns — Circuit breakers, quality gates and isolation — What does it mean?
The news. On October 1, 2026, Rudrendu Kumar Paul and Sourav Nandy posted Compound AI System Reliability: A Failure Taxonomy and Resilience Pattern Catalog from 150 Production Incidents (arXiv 2610.02503). They coded 97 incidents from public post-mortems and issue trackers of 12 open-source projects (LangChain, LlamaIndex, AutoGen, CrewAI, Haystack, DSPy and others) and 53 anonymized enterprise incidents into 23 failure modes in five categories: retrieval, generation, tool, orchestration and integration. Then they injected each failure mode into a six-component test system, 100 trials per scenario, with and without each fix. Read the paper →
Picture a restaurant kitchen. The pantry pulls ingredients, the stove cooks them, and a supplier delivers some items mid-service. Every station can be working by its own standard — the pantry hands over something, the stove produces a hot plate — and the diner still gets a bad meal. That is the paper's central argument: in a compound AI system, many of the failures that hurt users most live in the hand-offs between components, not inside any one of them. A retriever that returns the wrong documents still returns documents; a model fed the wrong documents still writes fluent text. Each health check measures its own station, so no alarm rings.
One enterprise incident in the paper is exactly this. A re-indexing job switched the document embedding model from text-embedding-ada-002 to text-embedding-3-small without updating the query encoder. Mean cosine similarity — the score of how closely each retrieved document matches the query — fell from 0.78 to 0.41, the retriever kept answering, and the generator produced confident answers from irrelevant documents for nine days until a user complained — a textbook case of the RAG failure modes that no single component reports. Across the corpus, silent failures took a mean of 4.2 days to detect, against 12 minutes for crashes.
Three guards for three ways a kitchen fails
The breaker deals with the late supplier. When one tool API slows down, the orchestrator keeps each request's context open while it waits, so a single slow dependency ties up memory and blocks other requests; in 8 of 10 timeout incidents, one slow tool cut system-wide throughput by 60–80% within minutes. A circuit breaker counts failures (the paper uses 5 failures in 30 seconds), then opens: callers get a structured fallback response at once instead of waiting for a timeout, and, by the usual design, it tries the component again after a cool-down (30 seconds in the paper). The AI-specific twist is what counts as a failure. Most AI failures return 200 OK with degraded content, so the breaker must watch quality signals such as retrieval relevance or output coherence, not only HTTP status codes. It is the same idea as a retry policy, run in reverse: decide when to stop trying.
The taster deals with the dish that looks right. An output quality gate is a cheap check between two stations. In the paper, one gate sits between retriever and generator and rejects a batch when the mean cosine similarity of the retrieved documents is below a calibrated threshold; a second gate checks the generator's answer for factual consistency against the retrieved context. This is the output filter idea moved from the edge of the system to the middle of it.
Separate burners deal with the station that hogs the kitchen. Component isolation gives each component its own memory, connection pool and rate-limit budget. When the tool executor burns through its rate-limit budget, it degrades alone instead of starving the retriever and the generator — a fail-safe rather than fail-open choice made at the level of resources.
Worked example: what each guard buys on the testbed
Hold three things fixed: the paper's six-component testbed, 1,000 concurrent requests (illustrative) when a fault hits, and the measured averages. These are separate measurements, not one combined run. Without a breaker, an injected fault reaches a mean of 3.8 components; with one, 0.4 — on average less than half a component, an 89% cut in cascade depth. Without isolation, a failure affects 78% of concurrent requests, which is 780 of the 1,000; with isolation it affects 28%, which is 280 — so 500 fewer requests are affected by the failure. Systems that run three or more of the patterns recover in a mean of 8.4 minutes instead of 28.7, saving about 20 minutes per incident. The price, also measured: a quality gate adds a median 120 ms per request, a breaker adds about 2 ms in normal operation but serves only fallbacks for its 30-second cool-down, and isolation raises peak memory by 18%.
| Pattern | What it measures | Effect on the testbed | Runtime cost | Source |
|---|---|---|---|---|
| Circuit breaker | cascade depth | 89% ±4% lower (3.8 → 0.4 components) | ~2 ms; 30 s of fallbacks when tripped | paper, Table 2 |
| Output quality gate | silent degradation caught | 73% ±6% caught before users | ~120 ms median per request | paper, Table 2 |
| Component isolation | blast radius | 64% ±5% lower (78% → 28% of requests) | ~18% more peak memory | paper, Table 2 |
| Semantic validator | semantic errors caught | 81% ±7% of errors that pass schema checks | ~80–150 ms | paper, Table 2 |
| Typed interfaces | integration failures | 92% ±3% removed | ~5–15 ms per boundary | paper, Table 2 |
Why one guard is never enough
No single pattern covers more than 40% of the failure modes, so the guards work as a set: the breaker contains cascades, the gate detects silent degradation, and isolation limits how many requests a failure reaches. The other two patterns fill specific holes. A semantic validator asks whether a tool's value is in the expected range or whether an answer addresses the original question. Typed interfaces replace loose JSON between components with checked schemas; in one incident, a JSON round trip between Python and TypeScript turned a threshold of 0.85 into 0.8500000000000001, a strict equality check failed, and 23% of results were silently dropped.
Read the numbers with the paper's own limits. They come from one six-component testbed under controlled fault injection, not live production; the authors call them indicative, not universal. The 71% recovery gain is measured against unstructured monitoring only — the authors note that a baseline with retries and exponential backoff would be stronger and leave it for future work. And 53 of the incidents come from a small number of organizations. For latency-sensitive systems, the paper suggests running quality gates asynchronously — log and alert on degradation without blocking the response — and keeping blocking gates for high-stakes outputs.
Goes deeper in: Agent Engineering → Production Harness Architecture → Why a Harness Fails in Production
Continue in trackProduction Harness Architecture — how a harness fails in productionRelated explainers
- ToolFailBench — tool-use failure taxonomy — a finer split of the tool category: skipped, ignored and fabricated tool results
- FakeLab — the fragmentation effect — another case where every per-component monitor passes while the system as a whole fails
- Chronicle — cut-point replay vs stubbed boundaries — turning an incident like these into a regression test