Program
3 sessions
Safety
-
Mechanistic Auditing of Deceptive Circuits
September 15, 2026 15:00 · Room 200
Interpretability tools for finding circuits that behave differently when a model thinks it is watched.
-
Scalable Oversight via Debate and Recursive Reward Models
September 16, 2026 14:00 · Aurora Hall
Supervising models on tasks where humans can no longer check the answer directly.
-
Robustness Certificates for Open-Weight Deployments
September 16, 2026 15:00 · Room 210
What guarantees you can and cannot make once model weights are public.