Compose multimodal datasets πΉ
β586Aug 15, 2026Updated this week
Alternatives and similar repositories for VQASynth
Users that are interested in VQASynth are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS'24] This repository is the implementation of "SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models"β337Dec 14, 2024Updated last year
- Official repo and evaluation implementation of VSI-Benchβ738Aug 5, 2025Updated last year
- The official repo for "SpatialBot: Precise Spatial Understanding with Vision Language Models.β349Jul 26, 2026Updated 3 weeks ago
- [NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligenceβ482Feb 5, 2026Updated 6 months ago
- SpaceR: The first MLLM empowered by SG-RLVR for video spatial reasoningβ111Jul 9, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [CVPR 2026] VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstructionβ440Jul 15, 2026Updated last month
- Awesome-LLM-3D: a curated list of Multi-modal Large Language Model in 3D world Resourcesβ2,249Apr 16, 2026Updated 4 months ago
- [NeurIPS'24] SpatialEval: a benchmark to evaluate spatial reasoning abilities of MLLMs and LLMsβ61Jan 23, 2025Updated last year
- [ICLR 2026] OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Modelsβ89Jan 21, 2026Updated 6 months ago
- Training recipe for SpatialReasoner [NeurIPS 2025]β45Apr 5, 2026Updated 4 months ago
- [ECCV2026] Visual Spatial Tuningβ201Mar 25, 2026Updated 4 months ago
- [ICCV 2025] A Simple yet Effective Pathway to Empowering LLaVA to Understand and Interact with 3D Worldβ388Oct 21, 2025Updated 9 months ago
- Synthetic VQA data generation code for SpatialReasoner.β20Nov 25, 2025Updated 8 months ago
- Official repository of Learning to Act from Actionless Videos through Dense Correspondences.β264Apr 25, 2024Updated 2 years ago
- End-to-end encrypted email - Proton Mail β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- [ICML 2024] 3D-VLA: A 3D Vision-Language-Action Generative World Modelβ630Oct 29, 2024Updated last year
- β12Jan 10, 2025Updated last year
- [CVPR 2026] Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Modelsβ178Feb 25, 2026Updated 5 months ago
- The code for paper 'Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors'β254Nov 28, 2025Updated 8 months ago
- Code for 3D-LLM: Injecting the 3D World into Large Language Modelsβ1,210Jun 6, 2024Updated 2 years ago
- [ICLR 2026] UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encodingβ63Jul 16, 2026Updated last month
- π₯ SpatialVLA: a spatial-enhanced vision-language-action model that is trained on 1.1 Million real robot episodes. Accepted at RSS 2025.β714Jun 23, 2025Updated last year
- Embodied Reasoning Question Answer (ERQA) Benchmarkβ287Mar 12, 2025Updated last year
- A Vision-Language Model for Spatial Affordance Prediction in Roboticsβ229Jul 17, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Evaluating and reproducing real-world robot manipulation policies (e.g., RT-1, RT-1-X, Octo) in simulation under common setups (e.g., Gooβ¦β1,139Dec 20, 2025Updated 7 months ago
- [NeurIPS 2025] Official implementation of "RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics"β266Dec 16, 2025Updated 8 months ago
- [ICML 2024] LEO: An Embodied Generalist Agent in 3D Worldβ488Apr 20, 2025Updated last year
- [CVPR 2026 Highlight] SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligenceβ84May 28, 2026Updated 2 months ago
- [NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modelingβ4,708Jun 26, 2026Updated last month
- [ICLR 2026 Oral (top 1.2%)] Official implementation of DepthLMβ368Jun 1, 2026Updated 2 months ago
- [TACL'23] VSR: A probing benchmark for spatial undersranding of vision-language models.β150Mar 25, 2023Updated 3 years ago
- A paper list for spatial reasoningβ774Jan 19, 2026Updated 6 months ago
- Reasoning in Space via Grounding in the World (ICLR 2025)β57Nov 3, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and clouβ¦β3,856Mar 12, 2026Updated 5 months ago
- STI-Bench : Are MLLMs Ready for Precise Spatial-Temporal World Understanding?β39Jan 12, 2026Updated 7 months ago
- [CVPR 2026] Scaling Spatial Intelligence with Multimodal Foundation Modelsβ294May 14, 2026Updated 3 months ago
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Modelsβ828Feb 20, 2025Updated last year
- β481Apr 14, 2026Updated 4 months ago
- From Flatland to Space (SPAR). Accepted to NeurIPS 2025 Datasets & Benchmarks. A large-scale dataset & benchmark for 3D spatial perceptioβ¦β95Jan 5, 2026Updated 7 months ago
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robotsβ1,650Aug 7, 2026Updated last week