Home /Research /Using Reinforcement Learning to Herd a Robotic Swarm to a Target\n Distribution
SWARM

Using Reinforcement Learning to Herd a Robotic Swarm to a Target\n Distribution

Zahi Kakish, Karthik Elamvazhuthi, Spring Berman

Year
2020
Citations
2
Access
Open access

Abstract

In this paper, we present a reinforcement learning approach to designing a\ncontrol policy for a "leader" agent that herds a swarm of "follower" agents,\nvia repulsive interactions, as quickly as possible to a target probability\ndistribution over a strongly connected graph. The leader control policy is a\nfunction of the swarm distribution, which evolves over time according to a\nmean-field model in the form of an ordinary difference equation. The dependence\nof the policy on agent populations at each graph vertex, rather than on\nindividual agent activity, simplifies the observations required by the leader\nand enables the control strategy to scale with the number of agents. Two\nTemporal-Difference learning algorithms, SARSA and Q-Learning, are used to\ngenerate the leader control policy based on the follower agent distribution and\nthe leader's location on the graph. A simulation environment corresponding to a\ngrid graph with 4 vertices was used to train and validate the control policies\nfor follower agent populations ranging from 10 to 100. Finally, the control\npolicies trained on 100 simulated agents were used to successfully redistribute\na physical swarm of 10 small robots to a target distribution among 4 spatial\nregions.\n

Keywords

Swarm behaviourReinforcement learningGraphComputer scienceVertex (graph theory)Mathematical optimizationRobotControl (management)GridSwarm robotics

Related papers

Browse all SWARM papers