Speaker
Dr. Julian Frost
Open Horizon AI
Speaks at
-
Mechanistic Auditing of Deceptive Circuits
September 15, 2026 15:00 · Room 200
Interpretability tools for finding circuits that behave differently when a model thinks it is watched.
-
Scalable Oversight via Debate and Recursive Reward Models
September 16, 2026 14:00 · Aurora Hall
Supervising models on tasks where humans can no longer check the answer directly.
Julian Frost leads interpretability and oversight research at Open Horizon AI. His recent work develops mechanistic tools for auditing model internals and scalable-oversight protocols for supervising systems that exceed human expertise.