Session

KV-Cache Compression at Long Context

Keeping the KV cache small enough to serve 200k-token contexts affordably.

Speakers

At long context the KV cache, not the weights, dominates memory. We compare eviction, quantization, and low-rank compression of the cache, and characterize where each preserves retrieval accuracy. A practical takeaway: a hybrid policy keeps recent and attended-to tokens at full precision while aggressively compressing the rest.