JMLR

Safe Learning Under Irreversible Dynamics via Asking for Help

Authors
Benjamin Plaut Juan Liévano-Karim Hanlin Zhu Stuart Russell
Paper Information
  • Journal:
    Journal of Machine Learning Research
  • Added to Tracker:
    Sep 08, 2026
Abstract

Most learning algorithms with formal regret guarantees essentially rely on trying all possible behaviors, which is problematic when some errors cannot be recovered from. Instead, we allow the learning agent to ask for help from a mentor and to transfer knowledge between similar states. We show that this combination enables the agent to learn both safely and effectively. Under standard online learning assumptions, we provide an algorithm whose regret and number of mentor queries are both sublinear in the time horizon for Markov decision processes with irreversible dynamics and infinite state spaces. Our proof involves a sequence of three reductions, making our result more general than a single algorithm. Conceptually, our result may be the first formal proof that it is possible for an agent to obtain high reward while becoming self-sufficient in an unknown, unbounded, and high-stakes environment without resets.

Author Details
Benjamin Plaut
Author
Juan Liévano-Karim
Author
Hanlin Zhu
Author
Stuart Russell
Author
Citation Information
APA Format
Benjamin Plaut , Juan Liévano-Karim , Hanlin Zhu & Stuart Russell . Safe Learning Under Irreversible Dynamics via Asking for Help. Journal of Machine Learning Research .
BibTeX Format
@article{paper1595,
  title = { Safe Learning Under Irreversible Dynamics via Asking for Help },
  author = { Benjamin Plaut and Juan Liévano-Karim and Hanlin Zhu and Stuart Russell },
  journal = { Journal of Machine Learning Research },
  url = { https://www.jmlr.org/papers/v27/25-2248.html }
}