Session

Data-Efficient Distillation from Mixture-of-Experts Teachers

Compressing a sparse expert model into a dense student without losing calibration.

Speakers

Distilling a large mixture-of-experts teacher into a compact dense student is complicated by the teacher's routed, uneven coverage. We present a distillation objective that matches expert-averaged logits and preserves the teacher's calibration, reaching strong student accuracy with a fraction of the usual distillation tokens.