Module preview

Regret minimization reframes strategic learning as a sequence of decisions judged against the best fixed choice in hindsight. We develop external, internal, and swap regret; regret matching and regret matching+; and no-regret dynamics whose average play converges to coarse correlated equilibrium, with a brief look at Blackwell's approachability. For LLM agents, the same framework models online tool selection and adaptive prompt choice as repeated no-regret decision problems—the bandit view of an agent deciding which tool to invoke.

Lectures and materials

Interactive demo