[NeurIPS'25] SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
☆40Oct 14, 2025Updated 9 months ago
Alternatives and similar repositories for SSR
Users that are interested in SSR are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACM MM'26] MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation☆64May 14, 2026Updated 2 months ago
- Score and Distribution Matching Policy: Advanced accelerated Visuomotor Policies via matched distillation☆11May 9, 2025Updated last year
- [ICLR 26] Context Tokens are Anchors: Understanding the Repeat Curse in dMLLMs from an Information Flow Perspective☆17Mar 6, 2026Updated 5 months ago
- [CVPR'25] Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception☆20Oct 11, 2025Updated 10 months ago
- VLA-RFT: Vision-Language-Action Models with Reinforcement Fine-Tuning☆161Oct 6, 2025Updated 10 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Training recipe for SpatialReasoner [NeurIPS 2025]☆45Apr 5, 2026Updated 4 months ago
- STI-Bench : Are MLLMs Ready for Precise Spatial-Temporal World Understanding?☆39Jan 12, 2026Updated 7 months ago
- [CVPR 2026]"Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World"☆17Jul 7, 2026Updated last month
- Learning 1D Causal Visual Representation with De-focus Attention Networks☆35Jun 7, 2024Updated 2 years ago
- [ICCV 2025] VLM4D: Towards Spatiotemporal Awareness in Vision Language Models☆55Nov 20, 2025Updated 8 months ago
- [ICLR 2025 Oral] Official Implementation for "Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Un…☆23Oct 24, 2024Updated last year
- Reshaping Action Error Distributions for Reliable Vision-Language-Action Models☆17Feb 5, 2026Updated 6 months ago
- ☆18Sep 10, 2025Updated 11 months ago
- Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal Transport☆15Feb 26, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ECCV2026] Visual Spatial Tuning☆201Mar 25, 2026Updated 4 months ago
- [ECCV'24 Oral] PiTe: Pixel-Temporal Alignment for Large Video-Language Model☆17Feb 13, 2025Updated last year
- VR-based Robot Teleoperation and Data Collection System for Humanoid Whole-Body VLA (Unitree G1)☆173Feb 17, 2026Updated 5 months ago
- Multi-Organ Foundation Model for Universal Ultrasound Image Segmentation with Task Prompt and Anatomical Prior☆15Sep 30, 2024Updated last year
- [CVPR 2026] HiF-VLA: An efficient, bidirectional spatiotemporal expansion Vision-Language-Action Model☆76Mar 11, 2026Updated 5 months ago
- ☆23May 8, 2025Updated last year
- Official implementation of Spatial-Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model [ICLR2026]☆281Jul 7, 2026Updated last month
- Findings of EMNLP 2023: InfoCL: Alleviating Catastrophic Forgetting in Continual Text Classification from An Information Theoretic Perspe…☆14Aug 13, 2024Updated 2 years ago
- [ICML 2026] Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models☆88May 18, 2026Updated 2 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- The code for paper 'Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors'☆253Nov 28, 2025Updated 8 months ago
- [CVPR 2026 Highlight] SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence☆84May 28, 2026Updated 2 months ago
- [ICLR 2026] Official implementation of the paper "📷 On the Generalization Capacities of MLLMs for Spatial Intelligence"☆29Mar 17, 2026Updated 4 months ago
- Official implementation for P2SAM (ACM MM 2024)☆14Dec 7, 2024Updated last year
- Code and datasets for "Text encoders are performance bottlenecks in contrastive vision-language models". Coming soon!☆11May 24, 2023Updated 3 years ago
- ViVa: A Video-Generative Value Model for Robot Reinforcement Learning☆85Jun 30, 2026Updated last month
- The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism☆31Jul 17, 2024Updated 2 years ago
- [NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence☆481Feb 5, 2026Updated 6 months ago
- Benchmarking Multi-Image Understanding in Vision and Language Models☆11Jul 29, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICCV2025] Official code repository of "CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction"☆62Aug 10, 2025Updated last year
- Official implementation of FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment☆55Mar 24, 2026Updated 4 months ago
- [ICCV 2025] Dynamic-VLM☆28Dec 16, 2024Updated last year
- [CVPR 2026] Official release of "Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning"☆135Apr 7, 2026Updated 4 months ago
- [ICLR 2026] SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models☆99Jun 9, 2026Updated 2 months ago
- ☆18Dec 25, 2025Updated 7 months ago
- ☆21May 7, 2024Updated 2 years ago