Reinforcement learning

Related papers: 20

About

Reinforcement learning (RL) is a machine learning paradigm in which an agent learns to make decisions by interacting with an environment, receiving numerical rewards or penalties based on its actions, and iteratively refining its behavior to maximize cumulative reward over time. Unlike supervised learning, RL requires no labeled training data; instead, the agent discovers effective strategies through trial and error guided by a reward signal. In robotics and AI, RL is used to train agents for complex tasks such as locomotion, dexterous manipulation, autonomous navigation, and game playing — often combining with deep neural networks (deep RL) to handle high-dimensional inputs like images and sensor streams. Multi-agent extensions enable coordinated or competitive behavior among multiple robots. RL matters because it provides a principled framework for automating the design of sophisticated control policies that are difficult or impossible to hand-engineer, enabling robots and AI systems to acquire adaptive, generalizable skills with minimal human specification of how tasks should be accomplished.

Top Cited Papers

Artificial intelligence: a modern approach

Citations: 22245 • 1995

A guide to deep learning in healthcare

Andre Esteva, Alexandre Robicquet, Bharath Ramsundar, Volodymyr Kuleshov, Mark A. DePristo, Katherine Chou, Claire Cui, Greg S. Corrado, Sebastian Thrun, Jeff Dean

Citations: 4608 • 2018

Deep Reinforcement Learning: A Brief Survey

Kai Arulkumaran, Marc Peter Deisenroth, Miles Brundage, Anil A. Bharath

Citations: 4261 • 2017

Trust Region Policy Optimization

John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, Pieter Abbeel

Citations: 3141 • 2015

Reinforcement learning in robotics: A survey

Jens Kober, J. Andrew Bagnell, Jan Peters

Citations: 3055 • 2013

A Comprehensive Survey of Multiagent Reinforcement Learning

Lucian Buşoniu, Robert Babuška, Bart De Schutter

Citations: 2178 • 2008

High-Dimensional Continuous Control Using Generalized Advantage Estimation

John Schulman, Philipp Moritz, Sergey Levine, Michael I. Jordan, Pieter Abbeel

Citations: 1750 • 2015

End-to-end training of deep visuomotor policies

Sergey Levine, Chelsea Finn, Trevor Darrell, Pieter Abbeel

Citations: 1715 • 2016

A State-of-the-Art Survey on Deep Learning Theory and Architectures

Md Zahangir Alom, Tarek M. Taha, Chris Yakopcic, Stefan Westberg, Paheding Sidike, Mst Shamima Nasrin, Mahmudul Hasan, Brian C. Van Essen, Abdul Ahad S. Awwal, Vijayan K. Asari

Citations: 1593 • 2019

Learning dexterous in-hand manipulation

OpenAI Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafał Józefowicz, Bob McGrew, Jakub Pachocki, Arthur J Petron, Matthias Plappert, Glenn Powell, Alex Ray, Jonas Schneider, Szymon Sidor, Josh Tobin, Peter Welinder, Lilian Weng, Wojciech Zaremba

Citations: 1588 • 2019

Target-driven visual navigation in indoor scenes using deep reinforcement learning

Yuke Zhu, Roozbeh Mottaghi, Eric Kolve, Joseph J. Lim, Abhinav Gupta, Li Fei-Fei, Ali Farhadi

Citations: 1507 • 2017

Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates

Shixiang Gu, Ethan Holly, Timothy Lillicrap, Sergey Levine

Citations: 1452 • 2017

End-to-End Training of Deep Visuomotor Policies

Sergey Levine, Chelsea Finn, Trevor Darrell, Pieter Abbeel

Citations: 1399 • 2015

Learning agile and dynamic motor skills for legged robots

Jemin Hwangbo, Joonho Lee, Alexey Dosovitskiy, C. Dario Bellicoso, Vassilios Tsounis, Vladlen Koltun, Marco Hutter

Citations: 1398 • 2019

Cooperative Multi-Agent Learning: The State of the Art

Liviu Panait, Sean Luke

Citations: 1250 • 2005

An Introduction to Deep Reinforcement Learning

Vincent François-Lavet, Peter Henderson, Riashat Islam, Marc G. Bellemare, Joëlle Pineau

Citations: 1246 • 2018

Multi-Agent Reinforcement Learning: A Selective Overview of Theories and Algorithms

Kaiqing Zhang, Zhuoran Yang, Tamer Başar

Citations: 1121 • 2021

A Survey of Actor-Critic Reinforcement Learning: Standard and Natural Policy Gradients

I. Grondman, Lucian Buşoniu, Gabriel A. D. Lopes, Robert Babuška

Citations: 1040 • 2012

Continuous control for robot based on deep reinforcement learning

Shansi Zhang

Citations: 937 • 2019

Reinforcement Learning and Dynamic Programming Using Function Approximators

Lucian Buşoniu, Robert Babuška, Bart De Schutter, Damien Ernst

Citations: 933 • 2010