ZhanYang-nwpu / Awesome-Multimodal-Large-Language-Models-for-UAV-Vision-Language-PerceptionView on GitHub
UAV-MLLMs
☆31Aug 23, 2026Updated last week
Alternatives and similar repositories for Awesome-Multimodal-Large-Language-Models-for-UAV-Vision-Language-Perception
Users that are interested in Awesome-Multimodal-Large-Language-Models-for-UAV-Vision-Language-Perception are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ISPRS'25] Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration☆18Jan 4, 2026Updated 7 months ago
- Parameter-Efficient Transfer Learning for Remote Sensing Image-Text Retrieval, 2023☆29Jan 14, 2024Updated 2 years ago
- ☆22Dec 2, 2025Updated 8 months ago
- Code for "MM-UAVBench: How Well Do Multimodal Large Language Models See, Think, and Plan in Low-Altitude UAV Scenarios?"☆23Jan 18, 2026Updated 7 months ago
- Multimodal Large Language Models for Remote Sensing (RS-MLLMs): A Survey☆402Aug 22, 2026Updated last week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- This is the open-sourced link of the TPAMI 2026 paper "SkyFind: A Large-Scale Benchmark Unveiling Referring Expression Comprehension for …☆33Aug 12, 2026Updated 2 weeks ago
- text-only training or language-free training for multimodal tasks (image/audio/video caption, retrieval, text2image)☆13Oct 15, 2024Updated last year
- RefDrone: A Challenging Benchmark for Drone Scene Referring Expression Comprehension☆47Jul 8, 2026Updated last month
- [AAAI'26] Code for the paper "AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning"☆18Jun 3, 2026Updated 2 months ago
- [ISPRS2025] SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model☆139Dec 1, 2025Updated 8 months ago
- Towards Open-Vocabulary Learing for Remote Sensing: A survey☆36Jul 5, 2026Updated last month
- [AAAI 2024] Mono3DVG: 3D Visual Grounding in Monocular Images, AAAI, 2024☆73Apr 9, 2024Updated 2 years ago
- [CVPR'25] ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval☆16Jun 17, 2026Updated 2 months ago
- Visual Object Tracking Paper List☆33Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆18Aug 22, 2024Updated 2 years ago
- Code for our paper "FocusTrack: A Self-Adaptive Local Sampling Algorithm for Efficient Anti-UAV Tracking"☆37Apr 21, 2025Updated last year
- 武大遥感院19级数据结构实习 包含1)CSV格式数据文件的读写 2)图的创建(邻接矩阵或邻接表) 3)图的遍历(广度优先或深度优先) 4)图的最短路径,并具体给出(A到B)的最短路径及其数值 5)最短路径的地图可视化展示 6)算法的时间和空间复杂度分析☆11Jan 24, 2021Updated 5 years ago
- [ACL'25 Oral] Code for the paper "UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban…☆31Aug 7, 2026Updated 3 weeks ago
- The HSI-MSI HAIHONG DATASET is used for for Hyperspectral and Multispectral Images Collaborative Classification.☆23Sep 11, 2024Updated last year
- [NeurIPS'25] FlySearch: Exploring how vision-language models explore