Session
Serving 100 Models on One GPU: Elastic Routing
Multiplexing a fleet of fine-tuned models on shared hardware with elastic routing.
Speakers
Teams increasingly serve dozens of fine-tuned variants of one base model. We describe an elastic routing layer that shares base weights, swaps adapters on demand, and reallocates GPU memory as traffic shifts, holding tail latency steady while packing a hundred models onto a single accelerator.