首页 /研究 /Dynamic Coverage Meets Regret: Unifying Two Control Performance Measures for Mobile Agents in Spatiotemporally Varying Environments
OTHER

Dynamic Coverage Meets Regret: Unifying Two Control Performance Measures for Mobile Agents in Spatiotemporally Varying Environments

Ben Haydon, Kirti D. Mishra, Patrick Keyantuo, Dimitra Panagou, F. K. Chow, Scott Moura, Chris Vermillion

发表年份
2021
引用次数
6

摘要

Numerous mobile robotic applications require agents to persistently explore and exploit spatiotemporally varying, partially observable environments. Ultimately, the mathematical notion of regret, which quite simply represents the instantaneous or time-averaged difference between the optimal reward and realized reward, serves as a meaningful measure of how well the agents have exploited the environment. However, while numerous theoretical regret bounds have been derived within the machine learning community, restrictions on the manner in which the environment evolves preclude their application to persistent missions. On the other hand, meaningful theoretical properties can be derived for the related concept of dynamic coverage, which serves as an exploration measurement but does not have an immediately intuitive connection with regret. In this paper, we demonstrate a clear correlation between an appropriately defined measure of dynamic coverage and regret, then go on to derive performance bounds on dynamic coverage as a function of the environmental parameters. We evaluate the correlation for several variants of an airborne wind energy system, for which the objective is to adjust the operating altitude in order to maximize power output in a spatiotemporally evolving wind field.

关键词

RegretComputer scienceMeasure (data warehouse)ExploitFunction (biology)Field (mathematics)Control (management)Artificial intelligenceMachine learningMathematics

相关论文

查看 OTHER 分类全部论文