Repository for awesome spatial/visual reasoning MLLMs. (focus more on embodied applications)
☆72Jun 26, 2025Updated last year
Alternatives and similar repositories for Awesome-spatial-visual-reasoning-MLLMs
Users that are interested in Awesome-spatial-visual-reasoning-MLLMs are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This is a project about visual spatial reasoning.☆146Jul 27, 2026Updated 3 weeks ago
- [NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence☆484Feb 5, 2026Updated 6 months ago
- A paper list for spatial reasoning☆774Jan 19, 2026Updated 7 months ago
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey☆1,018May 22, 2026Updated 2 months ago
- TorchHook: A PyTorch hooks manager, providing convenient interfaces to capture feature maps and debug models.☆18Oct 1, 2025Updated 10 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- 🎉Sonic IDE Desktop. Sonic IDE桌面版。☆100May 24, 2023Updated 3 years ago
- ☆12Jan 10, 2025Updated last year
- SpaceR: The first MLLM empowered by SG-RLVR for video spatial reasoning☆111Jul 9, 2025Updated last year
- Yuren 13B is an information synthesis large language model that has been continuously trained based on Llama 2 13B, which builds upon the…☆15Sep 25, 2023Updated 2 years ago
- [EMNLP 2024 Main] Official implementation of the paper "To Preserve or To Compress: An In-Depth Study of Connector Selection in Multimoda…☆16Dec 13, 2024Updated last year
- Latest open-source "Thinking with images" (O3/O4-mini) papers, covering training-free, SFT-based, and RL-enhanced methods for "fine-grain…☆115Aug 21, 2025Updated 11 months ago
- Collected the world's best computer vision labs and lecture materials.☆15Feb 23, 2025Updated last year
- Pixel-Level Reasoning Model trained with RL [NeuIPS25]☆304Jul 28, 2026Updated 3 weeks ago
- ECCV 2026 accepted CACFM code: RL-guided curvature-adaptive consistency flow matching for few-step FLUX/SDXL generation☆55Jun 25, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- 🔎Official code for our paper: "VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation".☆56Mar 18, 2025Updated last year
- 🚀 基于 Vue3、TypeScript、Vite 的企业级中后台快速开发框架,采用模块化设计,内置丰富的业务组件。☆26Oct 16, 2025Updated 10 months ago
- [ICLR 2025] Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision☆72Jul 10, 2024Updated 2 years ago
- This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-bas…☆1,441Aug 2, 2026Updated 2 weeks ago
- ☆27Jun 5, 2025Updated last year
- cooperative perception☆19Feb 6, 2026Updated 6 months ago
- Findings of EMNLP 2023: InfoCL: Alleviating Catastrophic Forgetting in Continual Text Classification from An Information Theoretic Perspe…☆14Aug 13, 2024Updated 2 years ago
- Official code release for paper "Robo-Imagine: A Robotic Video Generation Model, For Autoregressive Long-Term Task Video Generation With …☆31Jul 13, 2025Updated last year
- The development and future prospects of large multimodal reasoning models.☆614Jan 9, 2026Updated 7 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- A benchmark evaluates LLMs' performance in automating drawing revision tasks.☆60Updated this week
- ☆28May 14, 2025Updated last year
- [EMNLP 2025 Outstanding Paper Award] Official repo for DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph …☆22Nov 16, 2025Updated 9 months ago
- [ICRA 2023] Official implementation of the paper "ViPFormer: Efficient Vision-and-Pointcloud Transformer for Unsupervised Pointcloud Unde…☆24Mar 13, 2026Updated 5 months ago
- ☆18Mar 20, 2022Updated 4 years ago
- TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics☆22Nov 18, 2025Updated 9 months ago
- [CVPR2025] Official implementation of the paper "Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practi…☆48Oct 29, 2025Updated 9 months ago
- 这个算法用于无人机群避障一个加入机群的无人机,算法分为两种思路:(1)加入者的路径规划主动机动规避编队机群、(2)编队微调避让加入者。目前只做了第一种思路。唯一已知信息是原机群的运动轨迹F(x,y,z,t)|each plane,对于第一种思路:对于补位飞机唯一的输入参数是…☆43Aug 20, 2025Updated 11 months ago
- Official code for "Rethinking Chain-of-Thought Reasoning for Videos"☆21Dec 14, 2025Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [CVPR 2026 Fingdings] This repo is the official implementation of "Euclid’s Gift: Enhancing Spatial Perception and Reasoning in Vision‑La…☆28Mar 15, 2026Updated 5 months ago
- 🌍 A curated list of awesome AI agents for Remote Sensing applications 🛰️☆25Jul 3, 2026Updated last month
- [ICLR2025] γ -MOD: Mixture-of-Depth Adaptation for Multimodal Large Language Models☆45Oct 28, 2025Updated 9 months ago
- ☆12Sep 11, 2023Updated 2 years ago
- ☆16Jul 31, 2025Updated last year
- hydra-ai☆142May 8, 2025Updated last year
- Code release for VTW (AAAI 2025 Oral)☆68Nov 4, 2025Updated 9 months ago