Part IV · Decide and avoidChapter 12

Forty-four failure modes

The question How does context management fail, what is the tell for each failure, and how do you contain it?

8 min read Reference Step 12 of 20 2/4 in this part

In 30 seconds

  • Forty-four failures in eight clusters, each with a tell, a detector, a containment and a fix.
  • Use the symptom index during an incident; read the clusters to prevent the next one.
  • Four questions asked before a long task pre-empt most of the list.

You will be able to

  • Go from a symptom to its likely causes in under a minute
  • Contain a failure in-session and fix it permanently afterwards
  • Ask the four pre-emptive questions before a long task

Use this chapter two ways#

During an incident, start with the symptom index: find what you are seeing, then jump to the likely causes. Before a long task, read the four questions at the end. The underlying failure taxonomy is chapter 1's.

Every failure below has the same four fields:

  • Tell: what you notice.
  • Detect: how to confirm it.
  • Contain: what to do right now, in the session.
  • Fix: what to change so it does not happen again.

Forty-four failures in eight clusters

Eight clusters: dilution and displacement, staleness and poisoning, compaction, retrieval, isolation, cache and cost, session and state, and measurement. A · 7Dilution anddisplacementtoo much, too loud B · 6Staleness andpoisoningwrong, and trusted C · 7Compactionlost at the boundary D · 7Retrievalnever arrived E · 5Isolationsplit the wrong work F · 5Cache and costcheaper tokens, bigger bill G · 6Session andstatethe session outlived itself H · 1Measurementmeasured the wrong axis
Figure 1. Most real incidents combine two clusters: for example a compaction failure (C) that launders a stale fact (B). That is why the symptom index lists several causes per symptom.

Symptom-to-cause index#

What you see Likely causes, most likely first
The agent ignores a rule it followed earlier F-38, F-2, F-19
The agent repeats the same action F-39, F-18, F-14
Confidently wrong output, clean-looking transcript F-21, F-9, F-11, F-23
The agent asks for something already in context F-15, F-16, F-18
Costs rose after an "optimisation" F-33, F-35, F-36, F-37
The agent edits the wrong file F-23, F-10, F-8
The agent uses an odd tool F-1, F-24
The agent cannot tell it is finished F-16, F-41
Quality falls over a long session F-38, F-19, F-43, F-7
Parallel work produces incompatible pieces F-29, F-28
The bill tripled with no visible change F-37, F-33
The bug turned out to be in config F-24
It worked yesterday and not today F-44: variance you never measured

A: Dilution and displacement#

# Failure Tell Detect Contain Fix
F-1 Tool-definition bloat Prefix tax over 25K; capabilities never used Count schema tokens; count calls over 20 sessions Review the largest unused server; verify task coverage before disabling Tool audit; budget 20 tools, 15K tokens; re-audit quarterly
F-2 Instruction-file sprawl Over 200 lines; rules ignored late in sessions Line count over time; violations per rule Move the three most-violated rules to the end Inference test; scoped files; turn violated rules into hooks or lints
F-3 Tool-output flooding One command uses over 10% of the window Rank commands by token volume Re-run with a quiet flag; do not keep the log Output shaping; redirect to a file with a digest
F-4 Whole-file reading Read-utilisation under 5%; 1,200 lines read to change 4 Sample 20 reads Ask for a range or symbol Structural retrieval
F-5 Over-broad first search A grep returning 3,000+ hits, all loaded Result counts in transcripts Cap results; narrow before reading Search tools with hard caps; openers that name an identifier
F-6 Coherent-document distraction The agent cites the design doc for behaviour the code contradicts Trace wrong beliefs back to a document Remove the document; ask again Never pre-load design documents
F-7 Retention hoarding A 12K log from turn 8 still present at turn 60 Retention integral Offload or compact it Offload after a short delay

B: Staleness and poisoning#

# Failure Tell Detect Contain Fix
F-8 Stale self-read The agent cites a line its own edit moved Compare the in-context copy with disk Re-read the file Re-read before re-editing; auto-refresh after edits
F-9 Hallucination laundered by compaction A fact in the summary that never appeared in a tool result Find each summary fact's first appearance Reset to before the poison. Do not argue A "verified how" field in the schema; ground claims in tool output
F-10 Index staleness Semantic search returns pre-refactor code, no error Search for a symbol you just renamed Fall back to live search Prefer live retrieval; rebuild triggers
F-11 Stale memory entry The agent asserts a fact that was true six months ago Sample 20 entries against the repository Delete the entry Memories only about slow-changing things; prefer ADRs
F-12 Stale instruction rule A rule describing a refactor that already happened Read the file against the code quarterly Delete the line Review the instruction file in PRs that change what it describes
F-13 Prompt injection via retrieved content The agent follows an instruction found in a file, log or README Tool calls not traceable to a user turn Stop the session; audit what was retrieved Treat all retrieved content as data; never give tool output authority

C: Compaction#

# Failure Tell Detect Contain Fix
F-14 Compaction mid-sub-goal The agent restarts partly finished work Compare compaction times with sub-goal boundaries Restate the current sub-goal Semantic triggers with suppression rules
F-15 Detail loss The agent knows "there was an error" but not which Exact-string survival rate Re-run the command A verbatim exact-strings section
F-16 Lost sense of state The agent cannot tell whether it is done Termination recognition: 44.6% versus 77.2% in one study (AppWorld) S Ask "what is done and what remains?" A done / in-progress / not-started block; re-read the plan file
F-17 Post-compaction error spike The turn right after compaction fails more (+0.108 errors S) Error rate in the 3 turns after each boundary Insert a deliberate orientation step Make the plan-file re-read the first action
F-18 Regressive exploration Re-fetching and replaying right after a compaction Re-fetch rate in the next 10 turns Point at the offloaded artifact Offload before compacting
F-19 Compaction cascade Three or more compactions; a summary of summaries Compactions per session Reset now Cap at one; session discipline
F-20 Nondeterministic retention The same transcript compacted twice keeps different facts — model calls are stochastic D Run the compaction twice and diff Nothing in-session A strict schema; addressable stubs for observations

D: Retrieval#

# Failure Tell Detect Contain Fix
F-21 Starvation Confidently wrong output; clean transcript; no error Read-coverage audit: was the decisive file ever opened? Name the file and re-run Better retrieval quality, not more volume
F-22 Fragment blindness Reasoning as if code always runs when it sits inside a condition Check grep context settings Re-read the enclosing function At least -C 8; prefer symbol reads
F-23 Near-duplicate confusion The agent edits UserServiceV2 when UserService was live Grep for near-duplicate names Point at the right one Structural retrieval; better, delete the duplicate
F-24 Config blindness The agent exhausts the code and never checks YAML, env vars or flags Was the cause in a non-code file never searched? Search config explicitly Keep lexical search; include config globs by default
F-25 Silent zero-result Five irrelevant chunks returned; the agent proceeds as if it found something Relevance-sample 20 queries Switch to grep, which fails loudly Score thresholds with an explicit "no confident match"
F-26 Generated-code pollution Hits in node_modules/, dist/, __generated__/ Path distribution of search hits Re-run with exclusions Ignore files; exclusion defaults
F-27 Searching without narrowing Search turns over 60% of the session, no edits Search-to-edit turn ratio State the hypothesis and the file to check Better openers; expand along the reference graph

E: Isolation#

# Failure Tell Detect Contain Fix
F-28 Contract loss The parent re-derives something the sub-agent knew Redo rate after delegations Read the sub-agent's saved transcript A notable_outside_scope field; save transcripts
F-29 Conflicting implicit decisions Outputs individually coherent, mutually incompatible P Conflict rate across parallel delegations Reconcile by hand (expensive) The composability test; fix decisions in the brief
F-30 Prefix tax multiplication Total spend rises sharply with no reliability gain Total-token multiplier versus one agent Delegate coarser units Linear thread with disposable scouts
F-31 Verbose sub-agent results A transcript comes back instead of a result Parent growth per delegation Ask for the summary of the summary Strict output schema with a token cap
F-32 Unverifiable delegation Checking the output costs as much as producing it Time verifying versus doing it directly Stop delegating this kind of work Delegate only what is cheap to verify

F: Cache and cost#

# Failure Tell Detect Contain Fix
F-33 Cache thrash from a dynamic prefix Hit rate under 30%; cost per turn rising faster than tokens Cache hit rate per session Freeze the tool set and instruction file Prefix stability; static, hand-pruned tools
F-34 Timestamp in the prefix Near-zero hit rate despite a "stable" configuration Diff the first 2,000 tokens of two requests Remove the injected variable No timestamps, session IDs or counters in the prefix
F-35 Compaction costing more than it saves Cost rises after enabling compaction Cache-adjusted cost per turn, before and after Raise the threshold Compact less often; prefer resets
F-36 Optimising raw tokens A "20% token saving" that raised the bill Separate cached from uncached input Revert Report cache-adjusted cost
F-37 Sub-agent cost blindness The parent context looks great; the bill tripled Total spend across all agents Cap delegations per session Track the total-token multiplier

G: Session and state#

# Failure Tell Detect Contain Fix
F-38 Position decay of rules Rules followed early, ignored after two hours Violation rate by turn depth Restate the rule in the current turn Session caps; shorter instruction files; constraints next to the action
F-39 Distraction loop Same tool, same arguments, three or more times A loop detector Reset or redirect firmly; rephrasing does not work Reset triggers; a ruled-out list
F-40 Requirement clash Output satisfies a requirement withdrawn 20 turns ago Count requirement reversals Restate the full current requirement A living plan document as the single source of truth
F-41 Plan-file drift The plan says step 3; the agent is on step 6 Compare the plan with recent actions Rewrite the plan now Update at state changes; a "last updated at turn N" line
F-42 Handoff too thin The new session spends 20 turns re-establishing context Turns to first productive edit after a reset Paste in more detail A handoff with file:line pointers and verification commands
F-43 Sunk-cost session Hours into a degrading session, still pushing Quality against elapsed time Reset Mechanical caps: one compaction, or two hours

H: Measurement#

# Failure Tell Detect Contain Fix
F-44 Measuring the wrong axis "Compaction is basically free" from single runs, while users say "it worked yesterday" Compute Pass² alongside Pass@2; a widening gap is the signal S Nothing at runtime: it is a measurement defect k ≥ 2, always; track Pass² ÷ Pass@2

Four questions that pre-empt most of this#

Ask them before any long agent task.

  1. What is my prefix tax, and how much of it is dead? (F-1, F-2, F-33)
  2. What is the largest thing that will enter context, and where will it go afterwards? (F-3, F-7, F-18)
  3. If this session is compacted or reset, what must survive, and where is it written down? (F-14 to F-19, F-42)
  4. If this fails, will I be able to tell whether the agent was starved or diluted? (F-21, F-44)

Key takeaways

  1. Starvation can survive review because the trajectory may look clean.
  2. When a fact is poisoned, reset to before the poison; do not argue with it.
  3. Cost rising after an "optimisation" usually means the cache broke.
  4. "It worked yesterday" is a variance signal you never measured.

Terms used in this chapter

  • Retention integral — The sum of each segment's tokens multiplied by the number of turns it stays in context. The true cost of hoarding.
  • Poisoning — A failure mode where a false fact enters the context and is afterwards treated as established. It survives compaction and cannot be removed by appending a correction.
  • Laundering — Compaction turning a hedged hypothesis into an asserted fact by stripping the uncertainty around it.
  • Starvation — The agent lacks a fact, does not know it is missing, and proceeds on an assumption. A clean trajectory and a wrong answer.
  • Read-coverage — Whether the decisive file was ever opened during a failed task. It separates starvation from dilution, which need opposite fixes.
  • Fragment blindness — Reasoning about a code fragment without the control flow around it, such as a grep hit shown without the if statement that governs it.
  • Contract loss — A sub-agent knew something relevant but did not report it because its output contract did not ask.
  • Cache-adjusted cost — Input cost with cached and uncached tokens priced separately, about 0.10 versus 1.25 in relative units.
  • Position decay — A critical fact becoming harder to use because its position in a long input is less favourable.
  • Loop detector — A check that flags the same tool called with the same arguments three or more times with no new information in between.
  • Pass^k — Pass@k counts a task solved if any of k runs succeed ("can it ever?"). Pass^k, written Pass² for k = 2, requires all k runs to succeed ("can it reliably?").