Home /Research /Spiking Variational Policy Gradient for Brain Inspired Reinforcement Learning
LEARNING

Spiking Variational Policy Gradient for Brain Inspired Reinforcement Learning

Zhile Yang, Shangqi Guo, Ying Fang, Zhaofei Yu, Jian K. Liu

Year
2024
Citations
5

Abstract

Recent studies in reinforcement learning have explored brain-inspired function approximators and learning algorithms to simulate brain intelligence and adapt to neuromorphic hardware. Among these approaches, reward-modulated spike-timing-dependent plasticity (R-STDP) is biologically plausible and energy-efficient, but suffers from a gap between its local learning rules and the global learning objectives, which limits its performance and applicability. In this paper, we design a recurrent winner-take-all network and propose the spiking variational policy gradient (SVPG), a new R-STDP learning method derived theoretically from the global policy gradient. Specifically, the policy inference is derived from an energy-based policy function using mean-field inference, and the policy optimization is based on a last-step approximation of the global policy gradient. These fill the gap between the local learning rules and the global target. In experiments including a challenging ViZDoom vision-based navigation task and two realistic robot control tasks, SVPG successfully solves all the tasks. In addition, SVPG exhibits better inherent robustness to various kinds of input, network parameters, and environmental perturbations than compared methods.

Keywords

Reinforcement learningArtificial intelligenceComputer scienceMachine learningPattern recognition (psychology)

Related papers

Browse all LEARNING papers