Repeated-Game Convergence Laboratory
CSCE 631 · Module 3 (Regret Minimization, L7–L9). Two learners repeatedly play a small normal-form game under full-information, expected-payoff dynamics. Watch marginal averages, the joint empirical distribution, regret, and equilibrium gaps.
Strategies (current • average • trail)
Joint empirical play vs. product of marginals
The left matrix is the joint empirical distribution pT(a) = T−1∑t∏i xti,ai. CCE and CE are properties of this joint object — not of the product of marginals on the right. Their L1 distance measures temporal correlation.
L1 distance ||pT − x̄ȳT||1 = 0
Convergence diagnostics
Zero-sum game: the duality gap (sum exploitability) of the average strategies is shown; it should fall as averages converge.