Multimodal AI for UAV: Vision–Language Models in Human– Machine Collaboration
Maroš Krupáš, Ľubomír Urblík, Iveta Zolotová
- 发表年份
- 2025
- 引用次数
- 2
- 访问权限
- 开放获取
摘要
Recent advances in multimodal large language models (MLLMs)—particularly vision– language models (VLMs)—introduce new possibilities for integrating visual perception with natural-language understanding in human–machine collaboration (HMC). Unmanned aerial vehicles (UAVs) are increasingly deployed in dynamic environments, where adaptive autonomy and intuitive interaction are essential. Traditional UAV autonomy has relied mainly on visual perception or preprogrammed planning, offering limited adaptability and explainability. This study introduces a novel reference architecture, the multimodal AI–HMC system, based on which a dedicated UAV use case architecture was instantiated and experimentally validated in a controlled laboratory environment. The architecture integrates VLM-powered reasoning, real-time depth estimation, and natural-language interfaces, enabling UAVs to perform context-aware actions while providing transparent explanations. Unlike prior approaches, the system generates navigation commands while also communicating the underlying rationale and associated confidence levels, thereby enhancing situational awareness and fostering user trust. The architecture was implemented in a real-time UAV navigation platform and evaluated through laboratory trials. Quantitative results showed a 70% task success rate in single-obstacle navigation and 50% in a cluttered scenario, with safe obstacle avoidance at flight speeds of up to 0.6 m/s. Users approved 90% of the generated instructions and rated explanations as significantly clearer and more informative when confidence visualization was included. These findings demonstrate the novelty and feasibility of embedding VLMs into UAV systems, advancing explainable, human-centric autonomy and establishing a foundation for future multimodal AI applications in HMC, including robotics.
关键词
相关论文
Artificial intelligence: a modern approach
1995
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002
Are we ready for autonomous driving? The KITTI vision benchmark suite
Andreas Geiger, P Lenz, R. Urtasun
2012
TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems
Martı́n Abadi, Ashish Agarwal, Paul Barham 等 20 位作者
2016