Program
2 sessions
Inference
-
Speculative Decoding Beyond Draft Models
September 15, 2026 10:00 · Room 210
Getting speculative-decoding speedups without training and hosting a separate draft model.
-
KV-Cache Compression at Long Context
September 15, 2026 14:00 · Room 200
Keeping the KV cache small enough to serve 200k-token contexts affordably.