[NeurIPS'24] This repository is the implementation of "SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models"
β335Dec 14, 2024Updated last year
Alternatives and similar repositories for SpatialRGPT
Users that are interested in SpatialRGPT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Compose multimodal datasets πΉβ580Updated this week
- The official repo for "SpatialBot: Precise Spatial Understanding with Vision Language Models.β349Sep 14, 2025Updated 10 months ago
- Official repo and evaluation implementation of VSI-Benchβ732Aug 5, 2025Updated 11 months ago
- [NeurIPS'24] SpatialEval: a benchmark to evaluate spatial reasoning abilities of MLLMs and LLMsβ61Jan 23, 2025Updated last year
- [ICLR 2025 Oral] Official Implementation for "Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Unβ¦β22Oct 24, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- β12Jan 10, 2025Updated last year
- [ICLR 2026] OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Modelsβ88Jan 21, 2026Updated 6 months ago
- Synthetic VQA data generation code for SpatialReasoner.β20Nov 25, 2025Updated 7 months ago
- [NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligenceβ479Feb 5, 2026Updated 5 months ago
- [CVPR 2026] VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstructionβ428Updated this week
- A Vision-Language Model for Spatial Affordance Prediction in Roboticsβ227Jul 17, 2025Updated last year
- ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Groundingβ19Aug 8, 2025Updated 11 months ago
- [NeurIPS 2025] Official implementation of "RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics"β263Dec 16, 2025Updated 7 months ago
- SpaceR: The first MLLM empowered by SG-RLVR for video spatial reasoningβ111Jul 9, 2025Updated last year
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Training recipe for SpatialReasoner [NeurIPS 2025]β45Apr 5, 2026Updated 3 months ago
- [ECCV2026] Visual Spatial Tuningβ198Mar 25, 2026Updated 3 months ago
- The official repo for SpaceVista: All-Scale Visual Spatial Reasoning from mm to km.β43May 26, 2026Updated last month
- [ICCV 2025] A Simple yet Effective Pathway to Empowering LLaVA to Understand and Interact with 3D Worldβ384Oct 21, 2025Updated 9 months ago
- Orient Anything, ICML 2025β389Feb 6, 2026Updated 5 months ago
- Official code for paper: N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Modelsβ116Jan 14, 2026Updated 6 months ago
- [ICCV 2025] Official code for Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulationβ66Sep 12, 2025Updated 10 months ago
- [ICCV'25] Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awarenessβ70Jul 22, 2025Updated 11 months ago
- Github repository for "Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas" (ICML 2025)β76May 2, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [ICLR'26] This repository is the implementation of "3D Aware Region Prompted Vision Language Model"β28Feb 19, 2026Updated 5 months ago
- Official implementation of ECCV24 paper "SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding"β287Mar 19, 2025Updated last year
- SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imageryβ15Feb 1, 2026Updated 5 months ago
- The code for paper 'Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors'β245Nov 28, 2025Updated 7 months ago
- Awesome-LLM-3D: a curated list of Multi-modal Large Language Model in 3D world Resourcesβ2,238Apr 16, 2026Updated 3 months ago
- 3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understandingβ414Updated this week
- From Flatland to Space (SPAR). Accepted to NeurIPS 2025 Datasets & Benchmarks. A large-scale dataset & benchmark for 3D spatial perceptioβ¦β90Jan 5, 2026Updated 6 months ago
- β42Jun 9, 2025Updated last year
- (CVPR 26) Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Explorationβ35Mar 8, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official PyTorch implementation of CorrespondentDream: Enhancing 3D Fidelity of Text-to-3D using Cross-View Correspondences (CVPR 2024 Poβ¦β19Apr 29, 2024Updated 2 years ago
- [CVPR 2025] Source codes for the paper "3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning"β266Oct 2, 2025Updated 9 months ago
- Code release for 'Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs' (NeurIPS 2025)β31Oct 28, 2025Updated 8 months ago
- π₯ SpatialVLA: a spatial-enhanced vision-language-action model that is trained on 1.1 Million real robot episodes. Accepted at RSS 2025.β706Jun 23, 2025Updated last year
- THEORY OF SPACE: a benchmark for evaluating whether foundation models can actively explore under partial observability efficiently to buiβ¦β85Feb 27, 2026Updated 4 months ago
- [CVPR 2024] Probing the 3D Awareness of Visual Foundation Modelsβ354Dec 1, 2025Updated 7 months ago
- β17Oct 31, 2025Updated 8 months ago