A vision transformer with recurrent neural network-based fall activity recognition system for disabled persons in smart IoT environments
Abdulrahman Alzahrani, Asmaa Mansour Alghamdi
- Year
- 2025
- Citations
- 2
- Access
- Open access
Abstract
Falls are the primary basis of autonomy loss, injuries, and deaths among disabled persons and the elderly. With the development of technologies, falls are extensively researched by scientists to diminish severe consequences and adverse effects. Reliable fall recognition is crucial in humanoid robotics and healthcare research, as it helps minimize damage. The detection methods for falls are classified into three categories: wearable sensors, ambient sensors, and vision-based sensors. Over the last few years, computer vision and deep learning (DL) have been widely applied to fall detection systems. The application of DL for fall activity recognition has led to major developments in detection accuracy by overcoming numerous obstacles met by conventional models. This study presents a Vision Transformer and Self-Attention Mechanism with Recurrent Neural Network-Based Fall Activity Recognition System (VTSAMRNN-FARS) method. The primary objective of the VTSAMRNN-FARS method is to improve the fall detection and classification method for individuals with disabilities in smart IoT environments. Initially, the bilateral filtering (BF) model is used for image pre-processing to remove the noise in input image data. Furthermore, the feature extraction process is performed by the Vision Transformer (ViT) model to convert raw data into a reduced set of relevant features, thereby enhancing model performance and efficiency. For detecting fall activities, a bidirectional gated recurrent unit with a self-attention mechanism (BiGRU-SAM) model is implemented. Finally, the enhanced wombat optimization algorithm (EWOA) model optimally adjusts the hyperparameter values of the BiGRU-SAM approach, resulting in improved classification results. The simulation analysis of the VTSAMRNN-FARS methodology is examined under the UR_Fall_Dataset_Subset dataset. The comparison study of the VTSAMRNN-FARS methodology is reviewed and found to be 99.67% more effective than existing models.
Keywords
Related papers
Artificial intelligence: a modern approach
1995
Are we ready for autonomous driving? The KITTI vision benchmark suite
Andreas Geiger, P Lenz, R. Urtasun
2012
TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems
Martı́n Abadi, Ashish Agarwal, Paul Barham +17 more
2016
Vision meets robotics: The KITTI dataset
Andreas Geiger, Philip Lenz, Christoph Stiller +1 more
2013