Efficient video object segmentation based on frame-wise and segment-wise spatio-temporal interaction memory networks
Jisheng Dang, Huicheng ZHENG, Bimei Wang, Juncheng Li, Jianhuang Lai
- Year
- 2024
- Citations
- 3
Abstract
Video object segmentation aims to automatically segment objects of interest in videos, with wide applications in areas such as video editing, robot navigation, and autonomous driving. Existing methods for video object segmentation mostly rely on independent-frame appearance memory, which often falls short when dealing with complex video scenes with severe occlusions or appearance similarities. To address these challenges, this paper proposes a VOS method based on frame-wise and segment-wise spatio-temporal interaction memory (FSSTIM). FSSTIM introduces frame-wise and segment-wise spatio-temporal interaction memory construction blocks, which extract segment-wise spatio-temporal memory feature maps by constructing spatio-temporal context graph networks and enhance them by interacting with frame-wise memory feature maps, significantly improving the network's ability to handle similar appearances and object occlusions. Furthermore, the introduction of dynamic sampling memory readers achieves efficient multi-granularity historical information retrieval, speeding up inference and improving segmentation accuracy. Experiments on popular VOS datasets such as DAVIS, YouTube-VOS, and MOSE demonstrate that the proposed method achieves state-of-the-art performance while maintaining real-time processing speed and strong generalization capability.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991