Compose multimodal datasets πΉ
β589Aug 31, 2026Updated last week
Alternatives and similar repositories for VQASynth
Users that are interested in VQASynth are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS'24] This repository is the implementation of "SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models"β337Dec 14, 2024Updated last year
- Official repo and evaluation implementation of VSI-Benchβ739Aug 5, 2025Updated last year
- The official repo for "SpatialBot: Precise Spatial Understanding with Vision Language Models.β353Jul 26, 2026Updated last month
- [NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligenceβ486Feb 5, 2026Updated 7 months ago
- SpaceR: The first MLLM empowered by SG-RLVR for video spatial reasoningβ111Jul 9, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [CVPR 2026] VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstructionβ445Jul 15, 2026Updated last month
- Awesome-LLM-3D: a curated list of Multi-modal Large Language Model in 3D world Resourcesβ2,256Apr 16, 2026Updated 4 months ago
- [NeurIPS'24] SpatialEval: a benchmark to evaluate spatial reasoning abilities of MLLMs and LLMsβ61Jan 23, 2025Updated last year
- [ICLR 2026] OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Modelsβ94Jan 21, 2026Updated 7 months ago
- Training recipe for SpatialReasoner [NeurIPS 2025]β45Apr 5, 2026Updated 5 months ago
- [ECCV2026] Visual Spatial Tuningβ206Mar 25, 2026Updated 5 months ago
- [ICCV 2025] A Simple yet Effective Pathway to Empowering LLaVA to Understand and Interact with 3D Worldβ388Oct 21, 2025Updated 10 months ago
- Synthetic VQA data generation code for SpatialReasoner.β21Nov 25, 2025Updated 9 months ago
- Official repository of Learning to Act from Actionless Videos through Dense Correspondences.β263Apr 25, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [ICML 2024] 3D-VLA: A 3D Vision-Language-Action Generative World Modelβ634Oct 29, 2024Updated last year
- β12Jan 10, 2025Updated last year
- [CVPR 2026] Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Modelsβ179Feb 25, 2026Updated 6 months ago
- Code for 3D-LLM: Injecting the 3D World into Large Language Modelsβ1,214Jun 6, 2024Updated 2 years ago
- The code for paper 'Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors'β258Nov 28, 2025Updated 9 months ago
- [ICLR 2026] UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encodingβ63Jul 16, 2026Updated last month
- π₯ SpatialVLA: a spatial-enhanced vision-language-action model that is trained on 1.1 Million real robot episodes. Accepted at RSS 2025.β719Jun 23, 2025Updated last year
- Embodied Reasoning Question Answer (ERQA) Benchmarkβ290Mar 12, 2025Updated last year
- A Vision-Language Model for Spatial Affordance Prediction in Roboticsβ229Jul 17, 2025Updated last year
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Evaluating and reproducing real-world robot manipulation policies (e.g., RT-1, RT-1-X, Octo) in simulation under common setups (e.g., Gooβ¦β1,157Dec 20, 2025Updated 8 months ago
- [NeurIPS 2025] Official implementation of "RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics"β266Dec 16, 2025Updated 8 months ago
- [ICML 2024] LEO: An Embodied Generalist Agent in 3D Worldβ489Apr 20, 2025Updated last year
- [CVPR 2026 Highlight] SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligenceβ86May 28, 2026Updated 3 months ago
- [NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modelingβ4,730Jun 26, 2026Updated 2 months ago
- [ICLR 2026 Oral (top 1.2%)] Official implementation of DepthLMβ367Jun 1, 2026Updated 3 months ago
- [TACL'23] VSR: A probing benchmark for spatial undersranding of vision-language models.β150Mar 25, 2023Updated 3 years ago
- A paper list for spatial reasoningβ780Aug 23, 2026Updated 2 weeks ago
- Reasoning in Space via Grounding in the World (ICLR 2025)β57Nov 3, 2025Updated 10 months ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- STI-Bench : Are MLLMs Ready for Precise Spatial-Temporal World Understanding?β40Jan 12, 2026Updated 7 months ago
- [CVPR 2026] Scaling Spatial Intelligence with Multimodal Foundation Modelsβ299May 14, 2026Updated 3 months ago
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Modelsβ834Feb 20, 2025Updated last year
- β481Apr 14, 2026Updated 4 months ago
- From Flatland to Space (SPAR). Accepted to NeurIPS 2025 Datasets & Benchmarks. A large-scale dataset & benchmark for 3D spatial perceptioβ¦β95Jan 5, 2026Updated 8 months ago
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robotsβ1,709Updated this week
- Benchmark and training code for MindCube: spatial mental modeling in vision-language models from limited views.β169Aug 23, 2026Updated 2 weeks ago