Shentong Mo
Papers
1
Total Citations
4
H-Index
1
About
Shentong Mo is a rising researcher at the forefront of multimodal machine learning, with a focus on designing models that can process and integrate information from a diverse array of sensory inputs—far beyond the typical text and image pairs. In their influential work, "High-Modality Multimodal Transformer," Mo tackles the challenge of learning from many heterogeneous modalities, such as spoken language, gestures, force, and proprioception, which are common in robotics and human-computer interaction. This paper introduces a novel framework that quantifies both modality-specific and interaction heterogeneity, enabling more effective representation learning across numerous data types. While still early in their career, Mo's contributions are already gaining traction, with this work accumulating citations and laying a critical foundation for next-generation AI systems that must operate in rich, real-world environments. By pushing the boundaries of what multimodal transformers can handle, Mo is helping to pave the way for more perceptive and adaptable artificial intelligence.
Research Focus
Key Achievements
Top Papers
- 1