Fitted Q-iteration in Continuous Action-space MDPs

Abstract

We consider continuous state, continuous action batch reinforcement learning where the goal is to learn a good policy from a sufficiently rich trajectory generated by some policy. We study a variant of fitted Q-iteration, where the greedy action selection is replaced by searching for a policy in a restricted set of candidate policies by maximizing the average action values. We provide a rigorous analysis of this algorithm, proving what we believe is the first finite-time bound for value-function based algorithms for continuous state and action problems. Note: In retrospect, it would have been better to call this algorithm an actor-critic algorithm. The algorithm that we considers updates a policy and a value function (action-value function in this case).

BibTeX key: antos2007
entry type: inproceedings
booktitle: NIPS
year: 2007
pages: 9--16
crossref: NIPS20
pdf: papers/rlca.pdf
date-modified: 2010-11-25 00:50:53 -0700
date-added: 2010-08-28 17:38:14 -0600

BibSonomy

Fitted Q-iteration in Continuous Action-space MDPs

Abstract

Tags

Users

Comments and Reviewsshow / hide

Cite this publication

More citation styles

search on