Collection of the latest spatial, 3D, and video/temporal reasoning papers
☆37Sep 29, 2025Updated 11 months ago
Alternatives and similar repositories for awesome-spatial-reasoning
Users that are interested in awesome-spatial-reasoning are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Spatial Aptitude Training for Multimodal Langauge Models☆34Feb 8, 2026Updated 7 months ago
- SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery☆15Feb 1, 2026Updated 7 months ago
- When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning☆19Jun 2, 2026Updated 3 months ago
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation (ICLR 2026)☆22Apr 27, 2026Updated 4 months ago
- A paper list for spatial reasoning☆784Aug 23, 2026Updated 3 weeks ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆73Feb 4, 2026Updated 7 months ago
- Github repository for "Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas" (ICML 2025)☆75May 2, 2025Updated last year
- ☆27Jun 5, 2025Updated last year
- The official repo for ”[WACV2025] Towards Accurate Unified Anomaly Segmentation“☆16Apr 14, 2025Updated last year
- Official code for the paper "Adversarial Magnification to Deceive Deepfake Detection through Super Resolution"☆12Jun 26, 2023Updated 3 years ago
- [CVPR 2024] Tune-An-Ellipse: CLIP Has Potential to Find What You Want☆14Jan 5, 2025Updated last year
- To predict weekly games of the National Football League using game stats☆13Jun 13, 2020Updated 6 years ago
- ☆74Feb 12, 2026Updated 7 months ago
- SpaceR: The first MLLM empowered by SG-RLVR for video spatial reasoning☆111Jul 9, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [NeurIPS 2025] Official Implementation of paper "Sherlock: Self-Correcting Reasoning in Vision-Language Models"☆31Jun 4, 2026Updated 3 months ago
- A Python/C++ library for controlling the Franka Panda robot☆15Jan 17, 2024Updated 2 years ago
- Official repo for From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models☆34Nov 2, 2025Updated 10 months ago
- Benchmark and training code for MindCube: spatial mental modeling in vision-language models from limited views.☆172Aug 23, 2026Updated 3 weeks ago
- [ICLR 2026] Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and Vision☆237May 31, 2026Updated 3 months ago
- [Electronics'21] Facial Emotion Recognition Using Transfer Learning in the Deep CNN"☆22Jul 3, 2024Updated 2 years ago
- ☆14Apr 11, 2017Updated 9 years ago
- [CVPR 2026]SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning☆36Aug 4, 2026Updated last month
- Overlay gifs of memes on your video calls (works with Zoom, Google Meet, Microsoft Teams)☆11Apr 15, 2021Updated 5 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Official Implementation of "CLEFT: Language-Image Contrastive Learning with Efficient Large Language Model and Prompt Fine-Tuning" on MIC…☆18Feb 12, 2025Updated last year
- World Models That Know When They Don't Know: Controllable Video Generation with Calibrated Uncertainty☆26Sep 1, 2026Updated 2 weeks ago
- ☆19Apr 10, 2025Updated last year
- Lecture notes for a course on Decision and Game Theory for undergraduates studying AI☆13Dec 14, 2018Updated 7 years ago
- 一体化网页笔记批注、协作与专注辅助工具,同时提供个性化助学 Agent 赋能理解复习。☆20Sep 9, 2026Updated last week
- This repository collects and organises state‑of‑the‑art papers on spatial reasoning for Multimodal Vision–Language Models (MVLMs).☆324Feb 17, 2026Updated 7 months ago
- ☆10Nov 18, 2021Updated 4 years ago
- A Dead Simple and Modularized Multi-Modal Training and Finetune Framework. Compatible to any LLaVA/Flamingo/QwenVL/MiniGemini etc series …☆19Apr 24, 2024Updated 2 years ago
- [NeurIPS 2024] Efficient Large Multi-modal Models via Visual Context Compression☆66Feb 19, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Models and Codes for the paper Question Relevance in VQA: Identifying Non-Visual And False-Premise Questions☆14Aug 6, 2018Updated 8 years ago
- ☆18Jul 31, 2025Updated last year
- A Visualization Tool for GPU Occupancy on S Cluster.☆13Nov 16, 2022Updated 3 years ago
- Using CNN for reconstruction of images aberrated by atmospheric turbulence☆13Mar 15, 2019Updated 7 years ago
- Official repository for “Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space”☆18Jan 27, 2026Updated 7 months ago
- Official code for ICML 2024 paper, "Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language Models"☆19Jun 12, 2024Updated 2 years ago
- Official pytorch implementation of "RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language…☆14Dec 16, 2024Updated last year