Session
Curriculum Schedules for Trillion-Token Pretraining
How the order you feed data in changes what a trillion-token model ends up knowing.
Speakers
Data ordering is often treated as a nuisance parameter, but at trillion-token scale the curriculum schedule measurably shifts downstream capability and calibration. This talk presents a controlled study of staged mixtures, from broad web text to targeted high-quality domains, and shows where curriculum helps, where it hurts, and how to schedule it without a second full training run.