ZhanYang-nwpu / Awesome-Multimodal-Large-Language-Models-for-UAV-Vision-Language-PerceptionView on GitHub
UAV-MLLMs
☆30Apr 7, 2026Updated 4 months ago
Alternatives and similar repositories for Awesome-Multimodal-Large-Language-Models-for-UAV-Vision-Language-Perception
Users that are interested in Awesome-Multimodal-Large-Language-Models-for-UAV-Vision-Language-Perception are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ISPRS'25] Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration☆18Jan 4, 2026Updated 7 months ago
- ☆22Dec 2, 2025Updated 8 months ago
- Code for "MM-UAVBench: How Well Do Multimodal Large Language Models See, Think, and Plan in Low-Altitude UAV Scenarios?"☆22Jan 18, 2026Updated 6 months ago
- This is the open-sourced link of the TPAMI 2026 paper "SkyFind: A Large-Scale Benchmark Unveiling Referring Expression Comprehension for …☆33May 27, 2026Updated 2 months ago
- RefDrone: A Challenging Benchmark for Drone Scene Referring Expression Comprehension☆45Jul 8, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [AAAI'26] Code for the paper "AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning"☆18Jun 3, 2026Updated 2 months ago
- [ISPRS2025] SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model☆136Dec 1, 2025Updated 8 months ago
- [AAAI 2024] Mono3DVG: 3D Visual Grounding in Monocular Images, AAAI, 2024☆72Apr 9, 2024Updated 2 years ago
- [CVPR'25] ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval☆16Jun 17, 2026Updated last month
- Visual Object Tracking Paper List☆32Updated this week
- ☆18Aug 22, 2024Updated last year
- 武大遥感院19级数据结构实习 包含1)CSV格式数据文件的读写 2)图的创建(邻接矩阵或邻接表) 3)图的遍历(广度优先或深度优先) 4)图的最短路径,并具体给出(A到B)的最短路径及其数值 5)最短路径的地图可视化展示 6)算法的时间和空间复杂度分析☆11Jan 24, 2021Updated 5 years ago
- [ACL'25 Oral] Code for the paper "UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban…☆31Updated this week
- The HSI-MSI HAIHONG DATASET is used for for Hyperspectral and Multispectral Images Collaborative Classification.☆23Sep 11, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [NeurIPS'25] FlySearch: Exploring how vision-language models explore☆24Mar 12, 2026Updated 4 months ago
- Integration bluerov containing path planner, pid controller, and computer vision system☆12Dec 9, 2022Updated 3 years ago
- ☆17May 26, 2026Updated 2 months ago
- ☆32Nov 27, 2025Updated 8 months ago
- [CVPR 2026 Highlight] GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing☆48Jul 1, 2026Updated last month
- Code for https://arxiv.org/abs/2305.19643☆19Mar 14, 2024Updated 2 years ago
- YBTrack YOLOv5 + BYTE 无人机目标跟踪系统☆29Jan 31, 2024Updated 2 years ago
- [ACM MM 25] Official repo of "UEMM-Air: Enable UAVs to Undertake More Multi-modal Tasks"☆37Aug 20, 2025Updated 11 months ago
- ☆29Jun 9, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- PS2手柄☆13Dec 27, 2022Updated 3 years ago
- (ECCV2022) EAGAN: EAGAN: Efficient Two-stage Evolutionary Architecture Search for GANs☆12Sep 15, 2022Updated 3 years ago
- ☆54Dec 9, 2024Updated last year
- This repository contains the code for our paper: Enhancing Abnormality Grounding for Vision-Language Models with Knowledge Descriptions☆19Jun 24, 2025Updated last year
- [ISPRS2026] DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models☆32Mar 24, 2026Updated 4 months ago
- ☆10Aug 29, 2024Updated last year
- 基于全局注意力的改进YOLOv7-AC的水下场景目标检测系统☆16Dec 4, 2023Updated 2 years ago
- ☆504Mar 26, 2025Updated last year
- Awesome Remote Sensing Vision-Language Datasets☆96Jul 31, 2026Updated last week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- CAN bus on STM32F103C8T6 "Blue pill" uVision and CubeMX☆17Jan 1, 2019Updated 7 years ago
- Unlock the potential of latent diffusion models with MNIST! 🚀 Dive into reconstructing and generating digits using cutting-edge techniqu…☆16Jan 6, 2025Updated last year
- An offical repo for ECCV 2024 Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching☆119Jul 7, 2026Updated last month
- [ICCV 2023 oral] Official repository of the paper "Similarity Min-Max: Zero-Shot Day-Night Domain Adaptation"☆48Apr 15, 2024Updated 2 years ago
- PyTorch code for IEEE TCI2022 paper "Deep Hyperspectral Image Fusion Network with Iterative Spatio-Spectral Regularization"☆10May 10, 2022Updated 4 years ago
- ☆16Jun 3, 2026Updated 2 months ago
- ☆22Feb 23, 2026Updated 5 months ago