Program
3 sessions
Training
-
Curriculum Schedules for Trillion-Token Pretraining
September 15, 2026 09:00 · Aurora Hall
How the order you feed data in changes what a trillion-token model ends up knowing.
-
Gradient Noise as a Signal: Adaptive Batch Sizing
September 15, 2026 11:00 · Cobalt Theatre
Using the geometry of gradient noise to grow the batch size on the fly.
-
Data-Efficient Distillation from Mixture-of-Experts Teachers
September 16, 2026 09:00 · Aurora Hall
Compressing a sparse expert model into a dense student without losing calibration.