Compose multimodal datasets πΉ
β583Jul 28, 2026Updated this week
Alternatives and similar repositories for VQASynth
Users that are interested in VQASynth are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS'24] This repository is the implementation of "SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models"β336Dec 14, 2024Updated last year
- Official repo and evaluation implementation of VSI-Benchβ734Aug 5, 2025Updated 11 months ago
- The official repo for "SpatialBot: Precise Spatial Understanding with Vision Language Models.β349Updated this week
- [NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligenceβ480Feb 5, 2026Updated 5 months ago
- SpaceR: The first MLLM empowered by SG-RLVR for video spatial reasoningβ111Jul 9, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [CVPR 2026] VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstructionβ431Jul 15, 2026Updated 2 weeks ago
- Awesome-LLM-3D: a curated list of Multi-modal Large Language Model in 3D world Resourcesβ2,242Apr 16, 2026Updated 3 months ago
- [NeurIPS'24] SpatialEval: a benchmark to evaluate spatial reasoning abilities of MLLMs and LLMsβ61Jan 23, 2025Updated last year
- [ICLR 2026] OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Modelsβ89Jan 21, 2026Updated 6 months ago
- Training recipe for SpatialReasoner [NeurIPS 2025]β45Apr 5, 2026Updated 3 months ago
- [ECCV2026] Visual Spatial Tuningβ200Mar 25, 2026Updated 4 months ago
- [ICCV 2025] A Simple yet Effective Pathway to Empowering LLaVA to Understand and Interact with 3D Worldβ387Oct 21, 2025Updated 9 months ago
- Synthetic VQA data generation code for SpatialReasoner.β20Nov 25, 2025Updated 8 months ago
- Official repository of Learning to Act from Actionless Videos through Dense Correspondences.β262Apr 25, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ICML 2024] 3D-VLA: A 3D Vision-Language-Action Generative World Modelβ629Oct 29, 2024Updated last year
- β12Jan 10, 2025Updated last year
- [CVPR 2026] Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Modelsβ178Feb 25, 2026Updated 5 months ago
- The code for paper 'Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors'β248Nov 28, 2025Updated 8 months ago
- Code for 3D-LLM: Injecting the 3D World into Large Language Modelsβ1,211Jun 6, 2024Updated 2 years ago
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding. Accepted to ICLR 2026.β63Jul 16, 2026Updated last week
- π₯ SpatialVLA: a spatial-enhanced vision-language-action model that is trained on 1.1 Million real robot episodes. Accepted at RSS 2025.β710Jun 23, 2025Updated last year
- Embodied Reasoning Question Answer (ERQA) Benchmarkβ282Mar 12, 2025Updated last year
- A Vision-Language Model for Spatial Affordance Prediction in Roboticsβ227Jul 17, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Evaluating and reproducing real-world robot manipulation policies (e.g., RT-1, RT-1-X, Octo) in simulation under common setups (e.g., Gooβ¦β1,129Dec 20, 2025Updated 7 months ago
- [NeurIPS 2025] Official implementation of "RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics"β265Dec 16, 2025Updated 7 months ago
- [ICML 2024] LEO: An Embodied Generalist Agent in 3D Worldβ487Apr 20, 2025Updated last year
- [CVPR 2026 Highlight] SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligenceβ84May 28, 2026Updated 2 months ago
- [NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modelingβ4,630Jun 26, 2026Updated last month
- [ICLR 2026 Oral (top 1.2%)] Official implementation of DepthLMβ363Jun 1, 2026Updated last month
- [TACL'23] VSR: A probing benchmark for spatial undersranding of vision-language models.β149Mar 25, 2023Updated 3 years ago
- A paper list for spatial reasoningβ767Jan 19, 2026Updated 6 months ago
- Reasoning in Space via Grounding in the World (ICLR 2025)β56Nov 3, 2025Updated 8 months ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and clouβ¦β3,846Mar 12, 2026Updated 4 months ago
- STI-Bench : Are MLLMs Ready for Precise Spatial-Temporal World Understanding?β39Jan 12, 2026Updated 6 months ago
- [CVPR 2026] Scaling Spatial Intelligence with Multimodal Foundation Modelsβ290May 14, 2026Updated 2 months ago
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Modelsβ826Feb 20, 2025Updated last year
- β475Apr 14, 2026Updated 3 months ago
- From Flatland to Space (SPAR). Accepted to NeurIPS 2025 Datasets & Benchmarks. A large-scale dataset & benchmark for 3D spatial perceptioβ¦β90Jan 5, 2026Updated 6 months ago
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robotsβ1,585Jul 8, 2026Updated 3 weeks ago