Session

Long-Horizon Credit Assignment in Multi-Agent Teams

Figuring out which agent in a team deserves credit for an outcome many steps later.

Speakers

When a team of agents collaborates over a long horizon, assigning credit for a final outcome is hard. We compare counterfactual and value-decomposition approaches, and show that a learned per-agent advantage estimator improves both training stability and the interpretability of who did what.