Context Management
On this page
Three small vendored-and-split modules: core/budget.ts is read-only
token tracking; core/compaction.ts is the tail-preserving compaction
that kicks in once budget.ts reports over the hard limit;
core/llm-summarizer.ts is a real model call for summarizing what gets
compacted, not compaction.ts's own concatenate-and-cut default — kept
in its own file deliberately, since compaction.ts stays decoupled
from the rest of the codebase's own types.
How recovery actually runs
Compactor.recover() only acts once usage is over budgetTokens * softThreshold (default 0.7) — otherwise it returns messages
untouched. The most recent tailMessages (default 4) are never
drained or summarized, whatever else happens:
- Drain — cheap, no model call. Deletes from the oldest end of
what's left of
head, capped at half of it, so there's always something left for stage 2 to compress rather than deleting everything outright. If this alone gets back under target, it stops here (action: 'drained'). - Summarize — still over budget after draining: whatever remains
of
headgets passed to the configuredSummarizer, and the result replaces it as one message (action: 'summarized'). The tail stays exactly as it was.
The two Summarizer implementations
TruncatingSummarizer (default) |
LLMSummarizer |
|
|---|---|---|
| Model call | None | Yes — a real ModelCall, no tools |
| What it does | Concatenates the messages being compacted and cuts to maxChars (default 500) |
Prompts the model for a structured summary — facts established, actions taken, open question — verbatim where possible |
| Cost | Free, instant | One extra model call per compaction |
| Fidelity | Lossy — an early fact (an order id, a decision) can fall outside the cutoff and is gone for good | A fact stated once tends to survive, since the prompt calls it out as its own section |
| On failure | N/A — nothing to fail | Falls back to TruncatingSummarizer if the call throws or returns no usable text, rather than failing the whole turn |
When to use which
TruncatingSummarizer needs nothing wired up and is exactly what most
short-lived or test agents want — it exists so the whole recovery
pipeline is testable without an LLM at all. Reach for LLMSummarizer
once a conversation realistically runs long enough to compact more than
once and losing an early detail would actually break a later reply —
a support agent that needs to remember an order id from message 3 by
message 300, for instance. See Configure an Agent
for the actual AgentConfig.summarizer wiring, including pointing it
at a cheaper/faster model than the agent's own.
API-error recovery
core/recovery.ts is the adjacent, API-reliability half: the error
cases no provider SDK handles for you — a too-long request, an
oversized media attachment, a response truncated by hitting the output
token limit.