Session
Data-Efficient Distillation from Mixture-of-Experts Teachers
Compressing a sparse expert model into a dense student without losing calibration.
Speakers
Distilling a large mixture-of-experts teacher into a compact dense student is complicated by the teacher's routed, uneven coverage. We present a distillation objective that matches expert-averaged logits and preserves the teacher's calibration, reaching strong student accuracy with a fraction of the usual distillation tokens.