Home /Research /Temporal Object Detection in Videos Using Spatio-Temporal Transformers
PERCEPTION

Temporal Object Detection in Videos Using Spatio-Temporal Transformers

K. Shankar, K. Chandra Sekar, Renuka Devi S, B Monish, V Prathap, RVS Praveen

Year
2025
Citations
2

Abstract

Object detection in videos presents challenges such as motion blur, occlusions, and varying viewpoints, which traditional image-based models struggle to handle due to their inability to utilize temporal information. To overcome these limitations, this work introduces Spatio-Temporal Transformer-based Object Detection (ST-TOD), a novel framework that integrates spatial and temporal features for enhanced object detection and tracking in video sequences. The model combines Convolutional Neural Networks (CNNs) for spatial feature extraction with Transformer architectures to capture long-range dependencies across frames. A deformable attention mechanism dynamically focuses on key regions, mitigating motion blur and occlusions by emphasizing salient features. Experimental evaluations on benchmark video datasets demonstrate the effectiveness of ST-TOD, achieving 97 % accuracy, 87 % precision, and 90 % recall, significantly improving detection accuracy and robustness in complex environments. Its high efficiency makes it suitable for real-time applications such as video surveillance, autonomous driving, and robotics. By leveraging the strengths of spatial and temporal modeling, ST-TOD offers a powerful and adaptive solution for advanced video object detection tasks.

Keywords

Computer scienceArtificial intelligenceComputer visionTransformerObject detectionPattern recognition (psychology)EngineeringElectrical engineeringVoltage

Related papers

Browse all PERCEPTION papers