Article,

Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning?

S. Du, S. Kakade, R. Wang, and L. Yang.
(2019)cite arxiv:1910.03016.

Abstract

Modern deep learning methods provide an effective means to learn good representations. However, is a good representation itself sufficient for efficient reinforcement learning? This question is largely unexplored, and the extant body of literature mainly focuses on conditions which permit efficient reinforcement learning with little understanding of what are necessary conditions for efficient reinforcement learning. This work provides strong negative results for reinforcement learning methods with function approximation for which a good representation (feature extractor) is known to the agent, focusing on natural representational conditions relevant to value-based learning and policy-based learning. For value-based learning, we show that even if the agent has a highly accurate linear representation, the agent still needs to sample exponentially many trajectories in order to find a near-optimal policy. For policy-based learning, we show even if the agent's linear representation is capable of perfectly representing the optimal policy, the agent still needs to sample exponentially many trajectories in order to find a near-optimal policy. These lower bounds highlight the fact that having a good (value-based or policy-based) representation in and of itself is insufficient for efficient reinforcement learning. In particular, these results provide new insights into why the existing provably efficient reinforcement learning methods rely on further assumptions, which are often model-based in nature. Additionally, our lower bounds imply exponential separations in the sample complexity between 1) value-based learning with perfect representation and value-based learning with a good-but-not-perfect representation, 2) value-based learning and policy-based learning, 3) policy-based learning and supervised learning and 4) reinforcement learning and imitation learning.

BibTeX key: du2019representation
entry type: article
year: 2019
url: http://arxiv.org/abs/1910.03016
note: cite arxiv:1910.03016

Users

Comments and Reviewsshow / hide

Please log in to take part in the discussion (add own reviews or comments).

Cite this publication

%0 Journal Article %1 du2019representation %A Du, Simon S. %A Kakade, Sham M. %A Wang, Ruosong %A Yang, Lin F. %D 2019 %K generalization reinforcement-learning %T Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning? %U http://arxiv.org/abs/1910.03016 %X Modern deep learning methods provide an effective means to learn good representations. However, is a good representation itself sufficient for efficient reinforcement learning? This question is largely unexplored, and the extant body of literature mainly focuses on conditions which permit efficient reinforcement learning with little understanding of what are necessary conditions for efficient reinforcement learning. This work provides strong negative results for reinforcement learning methods with function approximation for which a good representation (feature extractor) is known to the agent, focusing on natural representational conditions relevant to value-based learning and policy-based learning. For value-based learning, we show that even if the agent has a highly accurate linear representation, the agent still needs to sample exponentially many trajectories in order to find a near-optimal policy. For policy-based learning, we show even if the agent's linear representation is capable of perfectly representing the optimal policy, the agent still needs to sample exponentially many trajectories in order to find a near-optimal policy. These lower bounds highlight the fact that having a good (value-based or policy-based) representation in and of itself is insufficient for efficient reinforcement learning. In particular, these results provide new insights into why the existing provably efficient reinforcement learning methods rely on further assumptions, which are often model-based in nature. Additionally, our lower bounds imply exponential separations in the sample complexity between 1) value-based learning with perfect representation and value-based learning with a good-but-not-perfect representation, 2) value-based learning and policy-based learning, 3) policy-based learning and supervised learning and 4) reinforcement learning and imitation learning.

@article{du2019representation, abstract = {Modern deep learning methods provide an effective means to learn good representations. However, is a good representation itself sufficient for efficient reinforcement learning? This question is largely unexplored, and the extant body of literature mainly focuses on conditions which permit efficient reinforcement learning with little understanding of what are necessary conditions for efficient reinforcement learning. This work provides strong negative results for reinforcement learning methods with function approximation for which a good representation (feature extractor) is known to the agent, focusing on natural representational conditions relevant to value-based learning and policy-based learning. For value-based learning, we show that even if the agent has a highly accurate linear representation, the agent still needs to sample exponentially many trajectories in order to find a near-optimal policy. For policy-based learning, we show even if the agent's linear representation is capable of perfectly representing the optimal policy, the agent still needs to sample exponentially many trajectories in order to find a near-optimal policy. These lower bounds highlight the fact that having a good (value-based or policy-based) representation in and of itself is insufficient for efficient reinforcement learning. In particular, these results provide new insights into why the existing provably efficient reinforcement learning methods rely on further assumptions, which are often model-based in nature. Additionally, our lower bounds imply exponential separations in the sample complexity between 1) value-based learning with perfect representation and value-based learning with a good-but-not-perfect representation, 2) value-based learning and policy-based learning, 3) policy-based learning and supervised learning and 4) reinforcement learning and imitation learning.}, added-at = {2019-10-10T19:53:17.000+0200}, author = {Du, Simon S. and Kakade, Sham M. and Wang, Ruosong and Yang, Lin F.}, biburl = {https://www.bibsonomy.org/bibtex/2ca1f707d8242f38d73430308ef09d0db/kirk86}, description = {[1910.03016] Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning?}, interhash = {3b5914b9680488c584ae49a61ed97805}, intrahash = {ca1f707d8242f38d73430308ef09d0db}, keywords = {generalization reinforcement-learning}, note = {cite arxiv:1910.03016}, timestamp = {2019-10-10T19:53:17.000+0200}, title = {Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning?}, url = {http://arxiv.org/abs/1910.03016}, year = 2019 }

BibSonomy

Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning?

Abstract

Tags

Users

Comments and Reviewsshow / hide

Cite this publication

More citation styles

search on