Session
Scalable Oversight via Debate and Recursive Reward Models
Supervising models on tasks where humans can no longer check the answer directly.
Speakers
How do you supervise a system on questions you cannot answer yourself? We compare debate, where two models argue and a judge decides, with recursive reward modeling, where oversight is bootstrapped level by level. We report where each protocol stays honest and where it degrades under optimization pressure.