Temporal Object Detection in Videos Using Spatio-Temporal Transformers
K. Shankar, K. Chandra Sekar, Renuka Devi S, B Monish, V Prathap, RVS Praveen
- Year
- 2025
- Citations
- 2
Abstract
Object detection in videos presents challenges such as motion blur, occlusions, and varying viewpoints, which traditional image-based models struggle to handle due to their inability to utilize temporal information. To overcome these limitations, this work introduces Spatio-Temporal Transformer-based Object Detection (ST-TOD), a novel framework that integrates spatial and temporal features for enhanced object detection and tracking in video sequences. The model combines Convolutional Neural Networks (CNNs) for spatial feature extraction with Transformer architectures to capture long-range dependencies across frames. A deformable attention mechanism dynamically focuses on key regions, mitigating motion blur and occlusions by emphasizing salient features. Experimental evaluations on benchmark video datasets demonstrate the effectiveness of ST-TOD, achieving 97 % accuracy, 87 % precision, and 90 % recall, significantly improving detection accuracy and robustness in complex environments. Its high efficiency makes it suitable for real-time applications such as video surveillance, autonomous driving, and robotics. By leveraging the strengths of spatial and temporal modeling, ST-TOD offers a powerful and adaptive solution for advanced video object detection tasks.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002