M3: Regret Minimization
Regret minimization reframes strategic learning as a sequence of decisions judged against the best fixed choice in hindsight. We develop external, internal, and swap regret; regret matching and regret matching+; and no-regret dynamics whose average play converges to coarse correlated equilibrium, with a brief look at Blackwell's approachability. For LLM agents, the same framework models online tool selection and adaptive prompt choice as repeated no-regret decision problems—the bandit view of an agent deciding which tool to invoke.
Module preview
Regret minimization reframes strategic learning as a sequence of decisions judged against the best fixed choice in hindsight. We develop external, internal, and swap regret; regret matching and regret matching+; and no-regret dynamics whose average play converges to coarse correlated equilibrium, with a brief look at Blackwell's approachability. For LLM agents, the same framework models online tool selection and adaptive prompt choice as repeated no-regret decision problems—the bandit view of an agent deciding which tool to invoke.
Lectures and materials
L7: Regret Minimization
External, internal, and swap regret; regret matching.
L8: Regret Minimization (cont.)
Regret matching+ and no-regret dynamics.
L9: Regret Minimization (cont.)
No-regret dynamics to coarse correlated equilibrium; Blackwell approachability.
Interactive demo
Repeated-Game Convergence Lab
Interactive simulation of no-regret learning dynamics whose average play converges to equilibrium in repeated games.