JMLR

Impatient Bandits: Optimizing for the Long-Term Without Delay

Authors
Kelly W. Zhang Thomas Baldwin-McDonald Kamil Ciosek Lucas Maystre Daniel Russo
Paper Information
  • Journal:
    Journal of Machine Learning Research
  • Added to Tracker:
    Sep 08, 2026
Abstract

Increasingly, recommender systems are tasked with improving users' long-term satisfaction. In this context, we study a content exploration task, which we formalize as a bandit problem with delayed rewards. There is an apparent trade-off in choosing the learning signal: waiting for the full reward to become available might take several weeks, slowing the rate of learning, whereas using short-term proxy rewards reflects the actual long-term goal only imperfectly. First, we develop a predictive model of delayed rewards that incorporates all information obtained to date. Rewards as well as shorter-term surrogate outcomes are combined through a Bayesian filter to obtain a probabilistic belief. Second, we devise a bandit algorithm that quickly learns to identify content aligned with long-term success using this new predictive model. We prove a regret bound for our algorithm that depends on the Value of Progressive Feedback, an information-theoretic metric that captures the quality of short-term leading indicators that are observed prior to the long-term reward. We apply our approach to a podcast recommendation problem, where we seek to recommend shows that users engage with repeatedly over two months. We empirically validate that our approach significantly outperforms methods that optimize for short-term proxies or rely solely on delayed rewards, as demonstrated by an A/B test in a recommendation system that serves hundreds of millions of users.

Author Details
Kelly W. Zhang
Author
Thomas Baldwin-McDonald
Author
Kamil Ciosek
Author
Lucas Maystre
Author
Daniel Russo
Author
Citation Information
APA Format
Kelly W. Zhang , Thomas Baldwin-McDonald , Kamil Ciosek , Lucas Maystre & Daniel Russo . Impatient Bandits: Optimizing for the Long-Term Without Delay. Journal of Machine Learning Research .
BibTeX Format
@article{paper1627,
  title = { Impatient Bandits: Optimizing for the Long-Term Without Delay },
  author = { Kelly W. Zhang and Thomas Baldwin-McDonald and Kamil Ciosek and Lucas Maystre and Daniel Russo },
  journal = { Journal of Machine Learning Research },
  url = { https://www.jmlr.org/papers/v27/25-0030.html }
}