Files
unprompted/drafts/2026-07-04-eviction.md
Stephen Donahue c6d7c8c211
Some checks failed
Deploy to S3 / deploy (push) Failing after 1m24s
Swap published columns for 24 entries from random.md
Retract the seven published columns into drafts/ and promote the 24 per-date
entries split out of Notes/random.md. Bodies are verbatim; titles, subjects
and slugs are new.

Make publication a matter of location. Notes/Notes.11tydata.js gives anything
under Notes/ permalink: false unless tagged `column`; drafts/drafts.11tydata.js
renders drafts at their real permalink under `yarn dev` and excludes them from
collections entirely in a build, so no template can surface one. Publishing is
now `git mv drafts/<file> columns/`, and moving it back retracts it.

Untagged notes were previously rendered into _site/ and deployed — including
Notes/random.md, which was live at /Notes/random/.
2026-08-02 22:45:19 +00:00

353 lines
19 KiB
Markdown

---
title: Eviction Notice
date: 2026-07-04
layout: column.njk
tags: [column, computing]
permalink: /columns/2026/eviction/
---
## Topics to hit
* session management in LLM chats
* Cache mgmt strategies in general
* securing sessions
---
## The detour: cache eviction, done properly
### The setup
A cache is a bounded fast store sitting in front of a larger, slower, or
more-expensive backing store. Registers cache RAM; RAM caches disk; a CDN
caches an origin; a browser caches the network. The bet is always the same:
**most requests will hit the small fast thing, so the average access looks fast
even though the truth lives somewhere slow.**
The bet only pays off if the right items are resident. And since the cache is
bounded, admitting something new usually means throwing something out. That
decision — *what to evict* — is the entire subject. Everything below is a
different guess at the same question: **which resident item am I least likely to
need soon (or cheapest to reconstruct if I'm wrong)?**
Hold onto the agent-session framing while you read. An LLM agent runs against a
fixed **context window** — a few hundred thousand tokens of working memory in
front of an effectively unbounded backing store (the full transcript, the
filesystem, tool outputs, the codebase, the web). The moment the session
outgrows the window, something has to leave. Context management *is* a cache
eviction policy over tokens, whether or not anyone designed it as one. So every
policy below is also a claim about how an agent should decide what to forget.
### The yardstick you can't have: Bélády's OPT
Bélády (1966) gave the provably optimal offline policy, usually called **OPT** or
**MIN**: when you must evict, evict the item whose *next* use is farthest in the
future.
$$\text{evict} \;=\; \arg\max_{i \,\in\, \text{cache}} \; \big(\text{next-use time of } i\big)$$
It's optimal for minimizing misses and completely unimplementable — it requires
knowing the future request stream. Its job is to be the ceiling: every real
policy is a heuristic trying to *predict* next-use distance from the past. Keep
OPT in mind for the agent case, because there we get a strange gift the CPU
never had — the model can partly *narrate its own future* ("I'm done with this
file"), which is a cheap approximation of clairvoyance. More on that later.
### The classics, and where each one breaks
**FIFO** — evict the oldest-inserted item. O(1), dead simple, and it ignores
usage entirely: a hot item gets evicted just for being old. It also suffers
**Bélády's anomaly** — giving it *more* cache can produce *more* misses.
FIFO is the naive "keep the last N tokens, drop the front" truncation, and it
fails in agents the same way it fails in hardware: it will happily discard the
original task statement or a hard constraint because they're old, not because
they're useless.
**LRU** — evict the Least Recently Used item. This is the workhorse. It bets on
**temporal locality**: recently touched ⇒ likely touched again soon. Classic
implementation is a hash map over a doubly linked list — move-to-front on
access, evict the tail, all O(1). LRU is a **stack algorithm**, so it's immune
to Bélády's anomaly. Its two real weaknesses:
- **No scan resistance.** One sweep over a large dataset (a full table scan, a
one-time file crawl) touches a flood of items exactly once and evicts the
entire genuine working set to make room for garbage you'll never see again.
- **Recency ≠ importance.** LRU assigns infinite value to a single recent touch.
In a conversation the task statement is rarely re-quoted verbatim yet is the
single most important thing in the window.
**LFU** — evict the Least Frequently Used item, by access count. Captures
popularity, which recency misses. Its failure modes are the mirror of LRU's:
- **Stale popularity / cache pollution.** An item that was hot last hour keeps
its high count forever and refuses to leave. Counts must **age** — an
exponential decay is typical:
$$\text{score}_i \leftarrow \text{score}_i \cdot e^{-\lambda \Delta t} + \mathbb{1}[\text{access}]$$
- **Cold start.** A brand-new item enters at count 1 and gets evicted
immediately, even if it's about to become the hottest thing in the cache
(the "one-hit wonder vs. future star" problem). Naive LFU is also O(log n)
on a heap — though there's a well-known O(1) construction with frequency
buckets (Shah, Matani, Mitzenmacher).
The whole rest of the field is a **recency-vs-frequency reconciliation**, because
neither signal dominates and the two mistakes above are opposite mistakes.
### Blends and adaptivity
**LRU-K** (O'Neil, O'Neil & Weikum, 1993) — remember the last *K* reference
times per item and evict on the *K*-th-most-recent one (the "backward
K-distance"). LRU-2 in particular cleanly separates items referenced twice from
one-hit wonders, folding a little frequency into recency. Costs more bookkeeping.
**2Q** (Johnson & Shasha, 1994) — a cheaper LRU-2 approximation. First-timers go
into a FIFO probation queue; only on a *second* reference do they graduate to the
main LRU. Scans wash through the probation queue and never pollute the main list
— scan resistance for almost free.
**LIRS** (Jiang & Zhang, 2002) — rank by **reuse distance** (inter-reference
recency) rather than plain recency, which is a much better next-use predictor.
Strongly scan-resistant; it's the ancestor of MySQL/InnoDB's buffer pool policy.
**CLOCK / second-chance** (Corbató, 1968) — an O(1) LRU approximation for OS page
replacement. Pages sit in a ring, each with a reference bit; a rotating hand
clears the bit or, if already clear, evicts. **CLOCK-Pro** (Jiang, Chen & Zhang,
2005) does for LIRS what CLOCK does for LRU. The lesson that matters downstream:
often you don't implement the ideal policy, you implement a cheap gadget that
*approximates* it well enough.
### ARC: stop tuning, start adapting
**ARC — Adaptive Replacement Cache** (Megiddo & Modha, IBM, 2003) is the elegant
one, and the closest hardware analog to what good agent context managers
actually do. It keeps four lists:
- **T1** — items seen once recently (the recency half)
- **T2** — items seen at least twice recently (the frequency half)
- **B1, B2** — *ghost* lists: the keys **recently evicted** from T1 and T2, with
the data thrown away but the metadata kept
A target parameter *p* splits capacity between T1 and T2. The trick is that ARC
learns *p* from its own recent mistakes: a hit in ghost list **B1** means "I
evicted a recency item I then wanted back" ⇒ grow T1; a hit in **B2** means the
same for frequency ⇒ grow T2. It continuously re-balances recency against
frequency with **no magic constants and no workload-specific tuning**, and it's
scan- and burst-resistant. (It's also patented by IBM, which is the boring
reason ZFS ships ARC while the Linux page cache doesn't — worth knowing.)
Two ideas from ARC get reused verbatim in the agent story below: **ghost
entries** (remember *that* something existed and was relevant even after you drop
its contents) and **self-tuning the recency/frequency split**.
### The move most people miss: admission, not just eviction
Every policy so far assumes the newcomer belongs in the cache and only argues
about the victim. **TinyLFU / W-TinyLFU** (Einziger, Friedman & Manes, 2017 —
the policy behind Java's Caffeine and Go's Ristretto) flips that:
$$\text{admit candidate } c \text{ over victim } v \iff \widehat{\text{freq}}(c) > \widehat{\text{freq}}(v)$$
On a miss, before evicting the chosen victim, estimate how popular the incoming
item actually is and **refuse to admit it if it's less popular than what it would
displace.** Frequencies are estimated cheaply and approximately with a
**Count-Min Sketch** that's periodically halved (aging), so the metadata is tiny.
W-TinyLFU fronts this with a small **window LRU** to catch fresh bursts, then
guards a larger SLRU main region with the TinyLFU admission filter. It posts
some of the best hit ratios known at a fraction of the metadata cost.
The reframing is the takeaway: **sometimes the correct action is to leave the
newcomer out entirely rather than evict a good resident to hold it.** For an
agent that means: not every verbose tool result has earned a place in the
window. Admission control is arguably the highest-leverage and most-neglected
lever in context engineering.
### When items aren't the same size or cost
CPU pages are uniform; web objects, files, and tool outputs are not. **GreedyDual-Size**
(Cao & Irani, 1997) generalizes recency to a cost/size utility. Each object gets
$$H(p) = L + \frac{c(p)}{s(p)}$$
where *c* is fetch cost, *s* is size, and *L* is a running "inflation" clock. You
evict the minimum-*H* object and set *L* to its *H* (so age accrues), and reset
*H(p)* on access. The instructive part is the ratio: **value per unit of space.**
That single idea — evict by value density, not by age — is exactly what a token
budget forces on an agent, where a 20-token hard constraint outweighs a
40,000-token file dump many times over.
### Concepts worth having names for
- **Temporal / spatial locality** — the empirical regularity every policy mines.
- **Working set** (Denning, 1968) — the set of items in active use over a window;
a cache earns its keep when it holds the working set and little else.
- **Reuse distance / stack distance** (Mattson et al., 1970) — the analytical
tool; the whole memory-hierarchy analysis rests on it.
- **Scan resistance** — robustness to one-shot floods.
- **Aging / decay** — frequency without forgetting is a slow-motion leak.
- **Admission vs. eviction** — the two knobs; most systems only turn one.
- **Ghost entries** — metadata about evicted items, kept to learn from misses.
- **Stack algorithms & Bélády's anomaly** — the theory floor: some policies are
monotone in cache size, some perversely aren't.
---
## A second cache, one level down (don't conflate these)
"Cache eviction in LLMs" names **two different things at two different levels**,
and it's worth nailing the distinction before the story, because they rhyme
loudly enough to blur:
1. **Context management (this column).** Message-level. *Outside* the model. You
decide which turns, tool results, and files occupy the prompt. The unit is a
semantic chunk; the policy is yours to write.
2. **KV-cache eviction.** Tensor-level. *Inside* the model. During
autoregressive decoding the attention keys/values for every past token are
cached so you don't recompute them; that store grows linearly with sequence
length and becomes the memory bottleneck. Policies here decide which *tokens'*
K/V tensors to drop:
- **StreamingLLM** (Xiao et al., 2023) — keep a few initial "**attention
sink**" tokens plus a sliding window; astonishingly, those first tokens
matter far past their content.
- **H2O** (Zhang et al., 2023) — evict tokens with low **accumulated attention**
("heavy hitters" stay).
- **Scissorhands** (Liu et al., 2023), **SnapKV** (Li et al., 2024),
**FastGen** (Ge et al., 2023) — variations on importance-scored token
eviction and per-head budgeting.
Same verb, different altitude. The neat part is that a single request is
governed by eviction policies at *both* levels simultaneously — you managing
messages, the runtime managing tensors — plus a **third** cache, the **prompt /
prefix KV cache** (Anthropic's prompt caching): a stable prompt *prefix* is
cached across calls, so reshuffling early context to save tokens can *invalidate*
that cache and cost you more than it saves. Eviction decisions at one level
perturb the economics at another. That interaction is a genuine, underexplored
design surface.
## Story tying it together
The throughline is the oldest idea in systems: the **memory hierarchy**.
Registers → L1 → L2 → L3 → RAM → SSD → disk → network, each tier a cache of the
one below it, each governed by a replacement policy, the whole stack engineered
so the fast tiers *usually* have what you need. MemGPT (Packer et al., 2023) made
the analogy explicit for LLMs: treat the **context window as RAM**, treat
external stores as **disk**, and let the model act as its own **memory-management
unit** — paging information in and out. Agent context management is, quite
literally, the memory hierarchy re-derived over semantic units. Which means
decades of replacement-policy research is sitting right there, reusable:
| Cache concept | Agent-context analog |
| --- | --- |
| Bounded fast store | The context window (finite tokens) |
| Backing store | Transcript on disk, files, vector DB, codebase, the web |
| Cache line / page | A message, turn, tool result, file chunk, memory note |
| Miss + penalty | Needed info no longer resident → re-read file, re-run tool, or **lost reasoning** (latency + tokens + sometimes unrecoverable) |
| Hit ratio | Fraction of needed info already in the window |
| FIFO / sliding window | Naive truncation ("keep the last N tokens") |
| LRU | Keep recently-referenced turns |
| LFU | Keep frequently-referenced facts |
| Pinning | System prompt, task goal, hard constraints — non-evictable |
| Ghost entries | Breadcrumbs to evicted content ("we discussed X, see file Y") |
| Admission control (TinyLFU) | Don't admit the whole 40k-token file; filter verbose tool output *before* it enters the window |
| Cost/size-aware (GreedyDual-Size) | Evict by **value per token**, not by age |
| Adaptive split (ARC) | Shift budget between recent-turns and reference-facts on the fly |
| Compressed cache (zswap/zram) | **Summarization / compaction** — lossy re-encode instead of drop |
| Tiering + fetch-on-miss (MemGPT) | External memory / RAG / scratchpad files; **recall = the fetch on miss** |
| OPT / Bélády | Evict what won't be needed before the task ends — and agents can *self-predict* this |
But the analogy isn't clean, and the places it *breaks* are where the
interesting research is. Four twists that make the agent version harder — and
more interesting — than the CPU version:
1. **Items are compressible.** A CPU cache line is opaque; you keep it or you
drop it. A conversation turn can be *summarized* — replaced by a lossy, smaller
version that still carries the gist. The only real hardware analog is the
**compressed-memory cache** (Linux zswap/zram: compress the page before
spilling it). Summarization/compaction is the single most-used agent technique
and it lives in a gap the classic policies barely cover.
2. **Value and size vary wildly and semantically.** Pages are uniform; tokens are
not. A one-line "never touch prod" constraint can outweigh a giant file dump.
This is GreedyDual-Size taken to its limit: **value density is everything**,
and value is semantic, not countable.
3. **Misses can be unrecoverable.** A CPU miss is just a slower fetch. Evicting a
chain of reasoning an agent can't reconstruct isn't a re-fetch — it's *gone*.
That raises the cost of a wrong eviction and argues hard for ghost breadcrumbs
and pinning over silent truncation.
4. **The cache steers its own workload.** This is the deep one. A CPU cache is
passive: the reference stream is fixed and the cache just tries to serve it.
An agent's context **determines what the agent does next**, which determines
the future access pattern — the workload is *endogenous*. The standard
"fixed reference string" analysis doesn't even apply. Evicting the wrong thing
doesn't just cause a miss; it can send the agent down a different path
entirely.
And one twist in our favor:
5. **Self-predicted reuse.** OPT is unattainable for a CPU because the future is
opaque. An agent can *tell you* when it's done with a file, or that a subtask
is closed. That's a cheap, imperfect approximation of Bélády's clairvoyance
from the inside — arguably the most promising lever context engineering has,
and it has no hardware analog at all.
The neat story, then: **agent context management is the memory hierarchy
reinvented over meaning instead of bytes** — same eviction question, same
recency/frequency tension, same admission-vs-eviction and pinning-and-ghosts
toolkit — but with items that can be *compressed* rather than dropped, *value*
that is semantic rather than counted, *misses* that can be fatal rather than
slow, a *workload the cache itself writes*, and a model that can *see enough of
its own future* to approach the optimum no CPU can reach.
*(Third topic — securing sessions — hangs off twist #3 and the ghost/external-memory
tier: what persists in context or spills to durable memory is an attack surface.
Prompt-injection content that survives a compaction, or poisons the recalled
memory store, is an eviction/admission problem wearing a security hat. Thread to
develop later.)*
## References
**Classic cache replacement**
- L. A. Bélády, "A study of replacement algorithms for a virtual-storage
computer," *IBM Systems Journal*, 1966. (OPT/MIN)
- P. J. Denning, "The working set model for program behavior," *CACM*, 1968.
- F. J. Corbató, "A paging experiment with the Multics system," 1968. (CLOCK)
- R. Mattson, J. Gecsei, D. Slutz, I. Traiger, "Evaluation techniques for storage
hierarchies," *IBM Systems Journal*, 1970. (stack distance, stack algorithms)
- E. O'Neil, P. O'Neil, G. Weikum, "The LRU-K page replacement algorithm for
database disk buffering," *SIGMOD*, 1993.
- T. Johnson, D. Shasha, "2Q: A low overhead high performance buffer management
replacement algorithm," *VLDB*, 1994.
- P. Cao, S. Irani, "Cost-aware WWW proxy caching algorithms," *USENIX USITS*,
1997. (GreedyDual-Size)
- S. Jiang, X. Zhang, "LIRS: An efficient low inter-reference recency set
replacement policy," *SIGMETRICS*, 2002.
- N. Megiddo, D. Modha, "ARC: A self-tuning, low overhead replacement cache,"
*USENIX FAST*, 2003 (and the *IEEE Computer* 2004 write-up).
- S. Jiang, F. Chen, X. Zhang, "CLOCK-Pro: An effective improvement of the CLOCK
replacement," *USENIX ATC*, 2005.
- K. Shah, A. Matani, M. Mitzenmacher, "An O(1) algorithm for implementing the
LFU cache eviction scheme," 2010.
- G. Einziger, R. Friedman, B. Manes, "TinyLFU: A highly efficient cache
admission policy," *ACM TOS*, 2017. (see also the Caffeine design notes)
**KV-cache eviction (tensor level)**
- G. Xiao et al., "Efficient streaming language models with attention sinks"
(StreamingLLM), 2023.
- Z. Zhang et al., "H2O: Heavy-hitter oracle for efficient generative inference
of large language models," *NeurIPS*, 2023.
- Z. Liu et al., "Scissorhands: Exploiting the persistence of importance
hypothesis for LLM KV cache compression," *NeurIPS*, 2023.
- Y. Li et al., "SnapKV: LLM knows what you are looking for before generation,"
*NeurIPS*, 2024.
- S. Ge et al., "Model tells you what to discard: Adaptive KV cache compression
for LLMs" (FastGen), *ICLR*, 2024.
**Agent context, memory, and long-context behavior**
- C. Packer et al., "MemGPT: Towards LLMs as operating systems," 2023. (the
memory-hierarchy analogy, now Letta)
- P. Lewis et al., "Retrieval-augmented generation for knowledge-intensive NLP
tasks," *NeurIPS*, 2020. (RAG = fetch-on-miss)
- N. Liu et al., "Lost in the middle: How language models use long contexts,"
*TACL*, 2023. (position bias — *where* in context you keep something matters)
- Anthropic, "Effective context engineering for AI agents" and the prompt-caching
/ context-editing docs. (admission, compaction, prefix caching in practice)