[NeurIPS'25] SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
☆41Oct 14, 2025Updated 10 months ago
Alternatives and similar repositories for SSR
Users that are interested in SSR are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACM MM'26] MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation☆65May 14, 2026Updated 3 months ago
- Score and Distribution Matching Policy: Advanced accelerated Visuomotor Policies via matched distillation☆11May 9, 2025Updated last year
- [ICLR 26] Context Tokens are Anchors: Understanding the Repeat Curse in dMLLMs from an Information Flow Perspective☆22Mar 6, 2026Updated 5 months ago
- [CVPR'25] Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception☆21Oct 11, 2025Updated 10 months ago
- VLA-RFT: Vision-Language-Action Models with Reinforcement Fine-Tuning☆160Oct 6, 2025Updated 10 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Training recipe for SpatialReasoner [NeurIPS 2025]☆45Apr 5, 2026Updated 4 months ago
- STI-Bench : Are MLLMs Ready for Precise Spatial-Temporal World Understanding?☆40Jan 12, 2026Updated 7 months ago
- [CVPR 2026]"Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World"☆17Jul 7, 2026Updated last month
- Learning 1D Causal Visual Representation with De-focus Attention Networks☆35Jun 7, 2024Updated 2 years ago
- [ICCV 2025] VLM4D: Towards Spatiotemporal Awareness in Vision Language Models☆55Nov 20, 2025Updated 9 months ago
- [ICLR 2025 Oral] Official Implementation for "Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Un…☆23Oct 24, 2024Updated last year
- OpenHelix: An Open-source Dual-System VLA Model for Robotic Manipulation☆395Aug 27, 2025Updated last year
- Official code for 'Learning to Rank for In-Context Example Retrieval'☆21Dec 20, 2025Updated 8 months ago
- Reshaping Action Error Distributions for Reliable Vision-Language-Action Models☆17Feb 5, 2026Updated 6 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆18Sep 10, 2025Updated 11 months ago
- [ECCV2026] Visual Spatial Tuning☆206Mar 25, 2026Updated 5 months ago
- [ECCV'24 Oral] PiTe: Pixel-Temporal Alignment for Large Video-Language Model☆17Feb 13, 2025Updated last year
- VR-based Robot Teleoperation and Data Collection System for Humanoid Whole-Body VLA (Unitree G1)☆175Feb 17, 2026Updated 6 months ago
- [CVPR 2026] HiF-VLA: An efficient, bidirectional spatiotemporal expansion Vision-Language-Action Model☆77Mar 11, 2026Updated 5 months ago
- Official implementation of Spatial-Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model [ICLR2026]☆281Jul 7, 2026Updated last month
- TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics☆23Nov 18, 2025Updated 9 months ago
- [CVPR 2026 Highlight] SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence☆86May 28, 2026Updated 3 months ago
- [ICML 2026] Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models☆92May 18, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ICLR 2026] Official implementation of the paper "📷 On the Generalization Capacities of MLLMs for Spatial Intelligence"☆30Mar 17, 2026Updated 5 months ago
- [ACL 2024] Multi-modal preference alignment remedies regression of visual instruction tuning on language model☆48Nov 10, 2024Updated last year
- Official implementation for P2SAM (ACM MM 2024)☆14Dec 7, 2024Updated last year
- [3DV 2025] Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model☆124May 30, 2025Updated last year
- Code and datasets for "Text encoders are performance bottlenecks in contrastive vision-language models". Coming soon!☆11May 24, 2023Updated 3 years ago
- ViVa: A Video-Generative Value Model for Robot Reinforcement Learning☆92Jun 30, 2026Updated 2 months ago
- The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism☆31Jul 17, 2024Updated 2 years ago
- [NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence☆485Feb 5, 2026Updated 6 months ago
- Benchmarking Multi-Image Understanding in Vision and Language Models☆11Jul 29, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official implementation of FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment☆55Mar 24, 2026Updated 5 months ago
- [ACM MM 2025] ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models☆18Jul 15, 2025Updated last year
- [ICCV 2025] Dynamic-VLM☆28Dec 16, 2024Updated last year
- [CVPR 2026] Official release of "Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning"☆138Apr 7, 2026Updated 4 months ago
- ☆71Aug 7, 2025Updated last year
- [ICLR 2026] SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models☆99Jun 9, 2026Updated 2 months ago
- ☆18Dec 25, 2025Updated 8 months ago