JMLR

Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features

Authors
Shangtong Zhang Jiuqi Wang
Research Topics
Time Series
Paper Information
  • Journal:
    Journal of Machine Learning Research
  • Added to Tracker:
    Jul 06, 2026
Abstract

Temporal difference (TD) learning with linear function approximation (linear TD) is a classic and powerful prediction algorithm in reinforcement learning. While it is well-understood that linear TD converges almost surely to a unique point, this convergence traditionally requires the assumption that the features used by the approximator are linearly independent. However, this linear independence assumption does not hold in many practical scenarios. This work is the first to establish the almost sure convergence of linear TD without requiring linearly independent features. We prove that the weight iterates of linear TD converge to a bounded set, and that the value estimates derived from the weights in that set are the same almost everywhere. We also establish a notion of local stability of the weight iterates. Importantly, we do not impose assumptions tailored to feature dependence and do not modify the linear TD algorithm. Key to our analysis is a novel characterization of bounded invariant sets of the mean ODE of linear TD.

Author Details
Shangtong Zhang
Author
Jiuqi Wang
Author
Research Topics & Keywords
Time Series
Research Area
Citation Information
APA Format
Shangtong Zhang & Jiuqi Wang . Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features. Journal of Machine Learning Research .
BibTeX Format
@article{paper1413,
  title = { Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features },
  author = { Shangtong Zhang and Jiuqi Wang },
  journal = { Journal of Machine Learning Research },
  url = { https://www.jmlr.org/papers/v27/24-1538.html }
}