Speaker
Rohan Bhatt
Vector Foundry
Speaks at
-
Serving 100 Models on One GPU: Elastic Routing
September 16, 2026 10:00 · Room 210
Multiplexing a fleet of fine-tuned models on shared hardware with elastic routing.
-
KV-Cache Compression at Long Context
September 15, 2026 14:00 · Room 200
Keeping the KV cache small enough to serve 200k-token contexts affordably.
Rohan Bhatt builds serving infrastructure at Vector Foundry. His work on elastic multi-model routing and KV-cache compression lets small teams serve large fleets of fine-tuned models on modest hardware.