Using Reinforcement Learning to Herd a Robotic Swarm to a Target\n Distribution
Zahi Kakish, Karthik Elamvazhuthi, Spring Berman
- Year
- 2020
- Citations
- 2
- Access
- Open access
Abstract
In this paper, we present a reinforcement learning approach to designing a\ncontrol policy for a "leader" agent that herds a swarm of "follower" agents,\nvia repulsive interactions, as quickly as possible to a target probability\ndistribution over a strongly connected graph. The leader control policy is a\nfunction of the swarm distribution, which evolves over time according to a\nmean-field model in the form of an ordinary difference equation. The dependence\nof the policy on agent populations at each graph vertex, rather than on\nindividual agent activity, simplifies the observations required by the leader\nand enables the control strategy to scale with the number of agents. Two\nTemporal-Difference learning algorithms, SARSA and Q-Learning, are used to\ngenerate the leader control policy based on the follower agent distribution and\nthe leader's location on the graph. A simulation environment corresponding to a\ngrid graph with 4 vertices was used to train and validate the control policies\nfor follower agent populations ranging from 10 to 100. Finally, the control\npolicies trained on 100 simulated agents were used to successfully redistribute\na physical swarm of 10 small robots to a target distribution among 4 spatial\nregions.\n
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002