[NeurIPS'24] This repository is the implementation of "SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models"
β338Dec 14, 2024Updated last year
Alternatives and similar repositories for SpatialRGPT
Users that are interested in SpatialRGPT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Compose multimodal datasets πΉβ590Updated this week
- Official repo and evaluation implementation of VSI-Benchβ743Aug 5, 2025Updated last year
- The official repo for "SpatialBot: Precise Spatial Understanding with Vision Language Models.β354Jul 26, 2026Updated last month
- [NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligenceβ489Feb 5, 2026Updated 7 months ago
- Synthetic VQA data generation code for SpatialReasoner.β21Nov 25, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [NeurIPS'24] SpatialEval: a benchmark to evaluate spatial reasoning abilities of MLLMs and LLMsβ61Jan 23, 2025Updated last year
- [CVPR 2026] VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstructionβ448Jul 15, 2026Updated 2 months ago
- Training recipe for SpatialReasoner [NeurIPS 2025]β45Apr 5, 2026Updated 5 months ago
- [ICLR 2026] OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Modelsβ94Jan 21, 2026Updated 7 months ago
- [ICLR 2025 Oral] Official Implementation for "Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Unβ¦β23Oct 24, 2024Updated last year
- SpaceR: The first MLLM empowered by SG-RLVR for video spatial reasoningβ111Jul 9, 2025Updated last year
- [NeurIPS 2025] Official implementation of "RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics"β266Dec 16, 2025Updated 9 months ago
- [ECCV2026] Visual Spatial Tuningβ209Mar 25, 2026Updated 5 months ago
- Official code for paper: N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Modelsβ119Jan 14, 2026Updated 8 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- From Flatland to Space (SPAR). Accepted to NeurIPS 2025 Datasets & Benchmarks. A large-scale dataset & benchmark for 3D spatial perceptioβ¦β96Jan 5, 2026Updated 8 months ago
- The code for paper 'Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors'β258Nov 28, 2025Updated 9 months ago
- 3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understandingβ415Jul 20, 2026Updated 2 months ago
- The official repo for SpaceVista: All-Scale Visual Spatial Reasoning from mm to km.β44May 26, 2026Updated 3 months ago
- [ICCV 2025] Official code for Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulationβ66Sep 12, 2025Updated last year
- Orient Anything, ICML 2025β408Feb 6, 2026Updated 7 months ago
- [ICLR 2026 Oral (top 1.2%)] Official implementation of DepthLMβ368Jun 1, 2026Updated 3 months ago
- Spatial Aptitude Training for Multimodal Langauge Modelsβ34Feb 8, 2026Updated 7 months ago
- [ICCV 2025] A Simple yet Effective Pathway to Empowering LLaVA to Understand and Interact with 3D Worldβ390Oct 21, 2025Updated 10 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- β12Jan 10, 2025Updated last year
- [CVPR 2025] The code for paper ''Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding''.β224Jun 4, 2025Updated last year
- A paper list for spatial reasoningβ784Aug 23, 2026Updated 3 weeks ago
- A Vision-Language Model for Spatial Affordance Prediction in Roboticsβ231Jul 17, 2025Updated last year
- [ICLR'26] This repository is the implementation of "3D Aware Region Prompted Vision Language Model"β31Feb 19, 2026Updated 7 months ago
- π₯ SpatialVLA: a spatial-enhanced vision-language-action model that is trained on 1.1 Million real robot episodes. Accepted at RSS 2025.β725Jun 23, 2025Updated last year
- [CVPR 2025] Source codes for the paper "3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning"β275Oct 2, 2025Updated 11 months ago
- Official implementation of ECCV24 paper "SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding"β291Mar 19, 2025Updated last year
- THEORY OF SPACE: a benchmark for evaluating whether foundation models can actively explore under partial observability efficiently to buiβ¦β87Feb 27, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [CVPR 2026] Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Modelsβ179Feb 25, 2026Updated 6 months ago
- Awesome-LLM-3D: a curated list of Multi-modal Large Language Model in 3D world Resourcesβ2,259Apr 16, 2026Updated 5 months ago
- β32Jun 24, 2024Updated 2 years ago
- [CVPR 2026] G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoningβ354Apr 18, 2026Updated 5 months ago
- SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imageryβ15Feb 1, 2026Updated 7 months ago
- [CVPR 2026 Highlight] SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligenceβ87May 28, 2026Updated 3 months ago
- Github repository for "Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas" (ICML 2025)β75May 2, 2025Updated last year