Compose multimodal datasets πΉ
β591Sep 21, 2026Updated this week
Alternatives and similar repositories for VQASynth
Users that are interested in VQASynth are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS'24] This repository is the implementation of "SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models"β338Dec 14, 2024Updated last year
- Official repo and evaluation implementation of VSI-Benchβ743Aug 5, 2025Updated last year
- The official repo for "SpatialBot: Precise Spatial Understanding with Vision Language Models.β355Jul 26, 2026Updated 2 months ago
- [NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligenceβ490Feb 5, 2026Updated 7 months ago
- SpaceR: The first MLLM empowered by SG-RLVR for video spatial reasoningβ111Jul 9, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [CVPR 2026] VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstructionβ449Jul 15, 2026Updated 2 months ago
- Awesome-LLM-3D: a curated list of Multi-modal Large Language Model in 3D world Resourcesβ2,263Apr 16, 2026Updated 5 months ago
- [NeurIPS'24] SpatialEval: a benchmark to evaluate spatial reasoning abilities of MLLMs and LLMsβ61Jan 23, 2025Updated last year
- [ICLR 2026] OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Modelsβ94Jan 21, 2026Updated 8 months ago
- Training recipe for SpatialReasoner [NeurIPS 2025]β45Apr 5, 2026Updated 5 months ago
- [ECCV2026] Visual Spatial Tuningβ211Mar 25, 2026Updated 6 months ago
- [ICCV 2025] A Simple yet Effective Pathway to Empowering LLaVA to Understand and Interact with 3D Worldβ390Oct 21, 2025Updated 11 months ago
- Synthetic VQA data generation code for SpatialReasoner.β21Nov 25, 2025Updated 10 months ago
- Official repository of Learning to Act from Actionless Videos through Dense Correspondences.β263Apr 25, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [ICML 2024] 3D-VLA: A 3D Vision-Language-Action Generative World Modelβ638Oct 29, 2024Updated last year
- β12Jan 10, 2025Updated last year
- [CVPR 2026] Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Modelsβ179Feb 25, 2026Updated 7 months ago
- Code for 3D-LLM: Injecting the 3D World into Large Language Modelsβ1,217Jun 6, 2024Updated 2 years ago
- The code for paper 'Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors'β259Nov 28, 2025Updated 10 months ago
- [ICLR 2026] UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encodingβ64Jul 16, 2026Updated 2 months ago
- π₯ SpatialVLA: a spatial-enhanced vision-language-action model that is trained on 1.1 Million real robot episodes. Accepted at RSS 2025.β727Jun 23, 2025Updated last year
- Embodied Reasoning Question Answer (ERQA) Benchmarkβ297Mar 12, 2025Updated last year
- A Vision-Language Model for Spatial Affordance Prediction in Roboticsβ231Jul 17, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available β’ AdRun AI, ML, and HPC workloads on powerful cloud GPUsβwithout limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Evaluating and reproducing real-world robot manipulation policies (e.g., RT-1, RT-1-X, Octo) in simulation under common setups (e.g., Gooβ¦β1,169Dec 20, 2025Updated 9 months ago
- [NeurIPS 2025] Official implementation of "RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics"β267Dec 16, 2025Updated 9 months ago
- [ICML 2024] LEO: An Embodied Generalist Agent in 3D Worldβ489Apr 20, 2025Updated last year
- [CVPR 2026 Highlight] SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligenceβ88May 28, 2026Updated 3 months ago
- [NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modelingβ4,744Jun 26, 2026Updated 3 months ago
- [ICLR 2026 Oral (top 1.2%)] Official implementation of DepthLMβ369Jun 1, 2026Updated 3 months ago
- [TACL'23] VSR: A probing benchmark for spatial undersranding of vision-language models.β150Mar 25, 2023Updated 3 years ago
- A paper list for spatial reasoningβ787Aug 23, 2026Updated last month
- Reasoning in Space via Grounding in the World (ICLR 2025)β57Nov 3, 2025Updated 10 months ago
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and clouβ¦β3,863Mar 12, 2026Updated 6 months ago
- STI-Bench : Are MLLMs Ready for Precise Spatial-Temporal World Understanding?β40Jan 12, 2026Updated 8 months ago
- [CVPR 2026] Scaling Spatial Intelligence with Multimodal Foundation Modelsβ306May 14, 2026Updated 4 months ago
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Modelsβ839Feb 20, 2025Updated last year
- β483Apr 14, 2026Updated 5 months ago
- From Flatland to Space (SPAR). Accepted to NeurIPS 2025 Datasets & Benchmarks. A large-scale dataset & benchmark for 3D spatial perceptioβ¦β97Jan 5, 2026Updated 8 months ago
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robotsβ1,762Updated this week