Home /Research /A Novel Visuo-Tactile Object Recognition Pipeline using Transformers with Feature Level Fusion
OTHER

A Novel Visuo-Tactile Object Recognition Pipeline using Transformers with Feature Level Fusion

John Doherty, Bryan Gardiner, Nazmul Siddique, Emmett Kerr

Year
2024
Citations
2

Abstract

The task of visuo-tactile object recognition is key in enabling robots to interact with humans and their environment in an efficient and effective manner. The differing statistical properties of visual images and tactile time-series data make visuo-tactile fusion non-trivial and complex. This work investigates the usage of Transformers to perform feature level fusion for visuo-tactile data, utilising the Transformer to generate temporal relationships between the visual and tactile data through its self-attention structure. The proposed pipeline is tested on the PHAC-2 dataset, and a complex ablation experiment is completed across a collection of leading activation functions. The proposed pipeline is demonstrated to achieve state-of-the-art accuracy for visuo-tactile object recognition on the PHAC-2 dataset, achieving a 94.3% accuracy when data from two tactile actions are considered.

Keywords

Computer scienceArtificial intelligenceComputer visionPipeline (software)FusionPattern recognition (psychology)Cognitive neuroscience of visual object recognitionTransformerFeature extractionFeature (linguistics)

Related papers

Browse all OTHER papers