STI-Bench : Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
☆40Jan 12, 2026Updated 7 months ago
Alternatives and similar repositories for STI-Bench
Users that are interested in STI-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SpaceR: The first MLLM empowered by SG-RLVR for video spatial reasoning☆111Jul 9, 2025Updated last year
- [CVPR 2026 Highlight] SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence☆86May 28, 2026Updated 3 months ago
- From Flatland to Space (SPAR). Accepted to NeurIPS 2025 Datasets & Benchmarks. A large-scale dataset & benchmark for 3D spatial perceptio…☆95Jan 5, 2026Updated 7 months ago
- ☆14Mar 9, 2024Updated 2 years ago
- Code release for 'Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs' (NeurIPS 2025)☆31Oct 28, 2025Updated 10 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Open studio for "Thinking with Spatial Code" (https://arxiv.org/pdf/2603.05591)☆21Mar 18, 2026Updated 5 months ago
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs☆71Mar 22, 2026Updated 5 months ago
- The official repo for SpaceVista: All-Scale Visual Spatial Reasoning from mm to km.☆44May 26, 2026Updated 3 months ago
- [ICLR 2026] OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models☆94Jan 21, 2026Updated 7 months ago
- ☆20Aug 7, 2025Updated last year
- ☆15Jan 7, 2026Updated 7 months ago
- [ICLR 2026] MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence☆108Apr 28, 2026Updated 4 months ago
- [ICML 2026] ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning☆90Jul 8, 2026Updated last month
- [NeurIPS'25] SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning☆41Oct 14, 2025Updated 10 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [NeurIPS 2025 Spotlight] Official implementation of Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence☆485Feb 5, 2026Updated 6 months ago
- [AAAI 2026 Oral] STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving Scenes☆17Jan 23, 2026Updated 7 months ago
- Code and dataset for paper "SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition"☆19Mar 17, 2026Updated 5 months ago
- [CVPR 2025] Program synthesis for 3D spatial reasoning☆63Jun 16, 2025Updated last year
- RoboFAC: A Comprehensive Framework for Robotic Failure Analysis and Correction☆45Jul 15, 2026Updated last month
- PyTorch implementation of the paper: CASAGPT: Cuboid Arrangement and Scene Assembly for Interior Design [CVPR 2025]☆15Apr 5, 2025Updated last year
- [CVPR 2026] VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction☆443Jul 15, 2026Updated last month
- MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence☆62Mar 11, 2026Updated 5 months ago
- ☆42Jun 9, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The code for paper 'Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors'☆256Nov 28, 2025Updated 9 months ago
- [ECCV 2026] ViewSpatial-Bench:Evaluating Multi-perspective Spatial Localization in Vision-Language Models☆83Mar 9, 2026Updated 5 months ago
- [EMNLP'26 Findings] OPD-Evolver☆43Jun 17, 2026Updated 2 months ago
- TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics☆23Nov 18, 2025Updated 9 months ago
- EPIC-Kitchens-100 Action Recognition baselines: TSN, TRN, TSM☆33Mar 15, 2022Updated 4 years ago
- ☆25Apr 12, 2025Updated last year
- A collection of algorithm pipelines for segmentation of aerial imagery implemented by PyTorch.☆10Jun 21, 2021Updated 5 years ago
- [CVPR 2025] Beacon3D: Object-centric Evaluation for 3D Grounding-QA☆28Nov 25, 2025Updated 9 months ago
- Benchmark and training code for MindCube: spatial mental modeling in vision-language models from limited views.☆168Aug 23, 2026Updated last week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Source code for EMNLP2022 paper "Finding Skill Neurons in Pre-trained Transformers via Prompt Tuning".☆19Mar 13, 2023Updated 3 years ago
- [NeurIPS 2025] Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing☆99Jul 27, 2025Updated last year
- Official Implementation of "Geometrically-Constrained Agent for Spatial Reasoning"☆92Apr 7, 2026Updated 4 months ago
- [ICLR'26] This repository is the implementation of "3D Aware Region Prompted Vision Language Model"☆31Feb 19, 2026Updated 6 months ago
- Compose multimodal datasets 🎹☆589Updated this week
- S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence☆87Jul 22, 2026Updated last month
- [Preprint] Any 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale☆58Apr 14, 2026Updated 4 months ago