首页 /研究 /Using Reinforcement Learning to Herd a Robotic Swarm to a Target\n Distribution
SWARM

Using Reinforcement Learning to Herd a Robotic Swarm to a Target\n Distribution

Zahi Kakish, Karthik Elamvazhuthi, Spring Berman

发表年份
2020
引用次数
2
访问权限
开放获取

摘要

In this paper, we present a reinforcement learning approach to designing a\ncontrol policy for a "leader" agent that herds a swarm of "follower" agents,\nvia repulsive interactions, as quickly as possible to a target probability\ndistribution over a strongly connected graph. The leader control policy is a\nfunction of the swarm distribution, which evolves over time according to a\nmean-field model in the form of an ordinary difference equation. The dependence\nof the policy on agent populations at each graph vertex, rather than on\nindividual agent activity, simplifies the observations required by the leader\nand enables the control strategy to scale with the number of agents. Two\nTemporal-Difference learning algorithms, SARSA and Q-Learning, are used to\ngenerate the leader control policy based on the follower agent distribution and\nthe leader's location on the graph. A simulation environment corresponding to a\ngrid graph with 4 vertices was used to train and validate the control policies\nfor follower agent populations ranging from 10 to 100. Finally, the control\npolicies trained on 100 simulated agents were used to successfully redistribute\na physical swarm of 10 small robots to a target distribution among 4 spatial\nregions.\n

关键词

Swarm behaviourReinforcement learningGraphComputer scienceVertex (graph theory)Mathematical optimizationRobotControl (management)GridSwarm robotics

相关论文

查看 SWARM 分类全部论文