[CVPR 2026]SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
β36Aug 4, 2026Updated last month
Alternatives and similar repositories for SpatialStack
Users that are interested in SpatialStack are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π Awesome lists of papers and codes about Large Vision-Language Modelsβ13Apr 1, 2024Updated 2 years ago
- [CVPR 2026] VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstructionβ445Jul 15, 2026Updated last month
- The code for paper 'Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors'β258Nov 28, 2025Updated 9 months ago
- SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imageryβ15Feb 1, 2026Updated 7 months ago
- This is a project on visual spatial reasoning tasks-SIBenchβ28Jan 12, 2026Updated 8 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- This is a collection of resources related with Time-series.β19May 21, 2024Updated 2 years ago
- ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation (CVPR'25)β21Apr 2, 2025Updated last year
- Official implementation of V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising (ECCV 2026)β27Jun 29, 2026Updated 2 months ago
- [ICCV 2025] Official code for Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulationβ66Sep 12, 2025Updated last year
- [NeurIPS 2025] Source codes for the paper "MindJourney: Test-Time Scaling with World Models for Spatial Reasoning"β152Nov 4, 2025Updated 10 months ago
- [ICCV 2025] VLM4D: Towards Spatiotemporal Awareness in Vision Language Modelsβ56Nov 20, 2025Updated 9 months ago
- [CVPR 2026] Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignmentβ28May 11, 2026Updated 4 months ago
- [CVPR 2026] HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Modelsβ42Jul 2, 2026Updated 2 months ago
- From Flatland to Space (SPAR). Accepted to NeurIPS 2025 Datasets & Benchmarks. A large-scale dataset & benchmark for 3D spatial perceptioβ¦β95Jan 5, 2026Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [NeurIPS 2025] Official Implementation of paper "Sherlock: Self-Correcting Reasoning in Vision-Language Models"β31Jun 4, 2026Updated 3 months ago
- Awesome lists of papers and codes about open-vocabulary perception, including both 3D and 2Dβ65Aug 25, 2026Updated 2 weeks ago
- Turbo Lossless - 1.33x Smaller, 2.93x Faster, Decode with 1 ADD operationβ15Apr 4, 2026Updated 5 months ago
- A paper list for spatial reasoningβ782Aug 23, 2026Updated 2 weeks ago
- β19Mar 10, 2026Updated 6 months ago
- β42Jun 9, 2025Updated last year
- Dai, J., Wang, A., Ni, B. et al. Facial albedo generation for 3D face reconstruction from a single image via a coarse-to-fine approach. Mβ¦β16Jul 16, 2026Updated last month
- List of papers wrote by Focoos AI research team!β12Jun 3, 2025Updated last year
- β18Apr 8, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- β13Jan 29, 2024Updated 2 years ago
- [ICRA 26] C^2ROPE: Causal Continuous Rotary Positional Encoding for 3D Large Multimodal-Models Reasoningβ28Feb 13, 2026Updated 7 months ago
- The official implementation of the paper "Large Scale Knowledge Washing"β10Jun 12, 2024Updated 2 years ago
- An open-source evaluation toolkit to evaluate MLLMs on Spatial Intelligence using the EASI protocolβ18Jul 1, 2026Updated 2 months ago
- β12Apr 16, 2022Updated 4 years ago
- Room classification network training and inference codeβ19Jun 26, 2024Updated 2 years ago
- [ICCV 2019] Monocular depth estimation from a single imageβ12May 20, 2022Updated 4 years ago
- Official repo and evaluation implementation of VSI-Benchβ740Aug 5, 2025Updated last year
- Code release for "ThermalNeRF: Thermal Radiance Fields"β18Aug 3, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- quality assessment for UIEβ14Dec 13, 2025Updated 8 months ago
- [AAAI 2025] Official data and code for "TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances"β15Sep 11, 2025Updated last year
- Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Modelsβ16Nov 4, 2023Updated 2 years ago
- Code for the paper "A Sea of Words: An In-Depth Analysis of Anchors for Text Data", AISTATS 2023β14Oct 26, 2024Updated last year
- VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMsβ18Feb 3, 2026Updated 7 months ago
- Official implementation of "What does CLIP know about a red circle? Visual Prompt Engineering for VLMs", ICCV 2023β12Sep 21, 2023Updated 2 years ago
- Official implementation of paper "Semantic Novelty Detection via Relational Reasoning"β16Jul 10, 2023Updated 3 years ago