Program
3 sessions
Inference
-
Speculative Decoding Beyond Draft Models
September 15, 2026 10:00 · Room 210
Getting speculative-decoding speedups without training and hosting a separate draft model.
-
KV-Cache Compression at Long Context
September 15, 2026 14:00 · Room 200
Keeping the KV cache small enough to serve 200k-token contexts affordably.
-
Serving 100 Models on One GPU: Elastic Routing
September 16, 2026 10:00 · Room 210
Multiplexing a fleet of fine-tuned models on shared hardware with elastic routing.