Mirror Descent and the Information Ratio

Tor Lattimore; Andras Gyorgy

Mirror Descent and the Information Ratio

Tor Lattimore , Andras Gyorgy

[Proceedings link] [PDF]

Session: Bandits, RL and Control 3 (A)

Session Chair: Ilja Kuzborskij

Poster: Poster Session 4

Abstract

Abstract: We establish a connection between the stability of mirror descent and the information ratio by Russo and Van Roy (2014). Our analysis shows that mirror descent with suitable loss estimators and exploratory distributions enjoys the same bound on the adversarial regret as the bounds on the Bayesian regret for information-directed sampling. Along the way, we develop the theory for information-directed sampling and provide an efficient algorithm for adversarial bandits for which the regret upper bound matches exactly the best known information-theoretic upper bound. Keywords: Bandits, partial monitoring, mirror descent, information theory.

Mirror Descent and the Information Ratio

Tor Lattimore , Andras Gyorgy

Summary presentation

Full presentation

Discussion