Off-policy Learning with Eligibility Traces: A Survey

In the framework of Markov Decision Processes, off-policy learning, that is the problem of learning a linear approximation of the value function of some fixed policy from one trajectory possibly generated by some other policy. We briefly review on-policy learning algorithms of the literature (gradient-based and least-squares-based), adopting a unified algorithmic view. Then, we highlight a systematic approach for adapting them to off-policy learning with eligibility traces. This leads to some known algorithms - off-policy LSTD(λ), LSPE(λ), TD(λ), TDC/GQ(λ) - and suggests new extensions - off-policy FPKF(λ), BRM(λ), gBRM(λ), GTD2(λ). We describe a comprehensive algorithmic derivation of all algorithms in a recursive and memory-efficent form, discuss their known convergence properties and illustrate their relative empirical behavior on Garnet problems. Our experiments suggest that the most standard algorithms on and off-policy LSTD(λ)/LSPE(λ) - and TD(λ) if the feature space dimension is too large for a least-squares approach - perform the best.

Data and Resources

Additional Info

Field Value
Source https://inria.hal.science/hal-00644516
Author Geist, Matthieu, Scherrer, Bruno
Maintainer CCSD
Last Updated May 11, 2026, 13:09 (UTC)
Created May 11, 2026, 13:09 (UTC)
Identifier hal-00644516
Language en
Rights https://about.hal.science/hal-authorisation-v1/
contributor SUPELEC-Campus Metz ; Ecole Supérieure d'Electricité - SUPELEC (FRANCE)
creator Geist, Matthieu
date 2013-05-11T00:00:00
harvest_object_id 09b200d5-474e-4ce7-ac80-0c2b6c5bd6ed
harvest_source_id 3374d638-d20b-4672-ba96-a23232d55657
harvest_source_title test moissonnage SELUNE
metadata_modified 2025-11-04T00:00:00
relation info:eu-repo/semantics/altIdentifier/arxiv/1304.3999
set_spec type:REPORT