首页 /研究 /Labeled Initialized Adaptive Play Q-learning for Stochastic Games
OTHER

Labeled Initialized Adaptive Play Q-learning for Stochastic Games

Andriy Burkov, Brahim Chaib-draa

发表年份
2007
引用次数
5

摘要

Recently, initial approximation of Q-values of the multiagent Q-learning by the optimal single-agent Q-values has shown good results in reducing the complexity of the learning process. In this paper, we continue in the same vein and give a brief description of the Initialized Adaptive Play Q-learning (IAPQ) algorithm while establishing an effective stopping criterion for this algorithm. To do that, we adapt a technique called “labeling” to the multiagent learning context. Our approach demonstrates good empirical behavior in multiagent coordination problems, such as two-robot grid world stochastic game. We show that our Labeled IAPQ (i) is able to converge faster than IAPQ by permitting a certain predefined value of learning error and (ii) it establishes an effective stopping criterion, which permits terminating the learning process at a near-optimal point with a flexible learning speed/quality tradeoff.

关键词

Q-learningComputer scienceContext (archaeology)Process (computing)Artificial intelligenceOptimal stoppingPoint (geometry)Mathematical optimizationAdaptive learningGrid

相关论文

查看 OTHER 分类全部论文