Official repo of From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
☆27Jun 23, 2026Updated 2 months ago
Alternatives and similar repositories for OSI-Bench
Users that are interested in OSI-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026 Spotlight] Code for miXed Discrete Diffusion Language Model☆31Mar 16, 2026Updated 6 months ago
- [NeurIPS 2024] Artemis: Towards Referential Understanding in Complex Videos☆27Apr 8, 2025Updated last year
- [AAAI2025] ChatterBox: Multi-round Multimodal Referring and Grounding, Multimodal, Multi-round dialogues☆62May 2, 2025Updated last year
- Implementation of paper "CC-Diff: Enhancing Contextual Coherence in Remote Sensing Image Synthesis"☆28Dec 19, 2025Updated 9 months ago
- [CVPR 2026] LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding☆56Jul 7, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICCV 2023] Generative Prompt Model for Weakly Supervised Object Localization☆57Nov 10, 2023Updated 2 years ago
- Office implementation of "3D-Aware Hypothesis & Verification for Generalizable Relative Object Pose Estimation", ICLR 2024☆12Nov 5, 2024Updated last year
- [CVPR 2025] Adaptive Keyframe Sampling for Long Video Understanding☆237Dec 19, 2025Updated 9 months ago
- (CVPR2023/TPAMI2024) Integrally Pre-Trained Transformer Pyramid Networks -- A Hierarchical Vision Transformer for Masked Image Modeling☆218Jul 28, 2024Updated 2 years ago
- ☆77Mar 1, 2023Updated 3 years ago
- [ECCV 2024] ControlCap: Controllable Region-level Captioning☆81Oct 25, 2024Updated last year
- ☆31Sep 24, 2024Updated last year
- [CVPR 2026] Official repository of "StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synth…☆18Feb 21, 2026Updated 6 months ago
- [ICLR 2026] Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization☆32Mar 6, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- CVPR 2026 (Highlight); Spatial Intelligence; MLLMs☆48Feb 24, 2026Updated 6 months ago
- Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation.☆16Jul 21, 2025Updated last year
- Group project "Algorithms for large-scale optimal transport". Implement ADMMs and Sinkhorn's Algorithms.☆11Jan 28, 2019Updated 7 years ago
- Code for the paper "Data Attribution for Text-to-Image Models by Unlearning Synthesized Images."☆17May 23, 2025Updated last year
- Holistic Evaluation of Multimodal LLMs on Spatial Intelligence☆128Jul 1, 2026Updated 2 months ago
- Official PyTorch Implementation of "Better Source, Better Flow: Learning Condition-Dependent Source Distribution for Flow Matching"☆39Mar 1, 2026Updated 6 months ago
- Official code release of "DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding" [ICCV2025 Highlight]☆54Sep 27, 2025Updated 11 months ago
- ☆14Jan 4, 2025Updated last year
- ☆131Aug 27, 2026Updated 3 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- The repository of the ACCV 2024 paper "FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Ge…☆12Aug 15, 2026Updated last month
- The public reproducible analysis code used for the gaze project☆11May 16, 2026Updated 4 months ago
- Code repository for MMUGL: Multi-modal Graph Learning over UMLS Knowledge Graphs☆11Dec 7, 2023Updated 2 years ago
- An implementation of unsupervised example of the Forward-Forward algorithm proposed by (Hinton, 2022)☆10Jun 19, 2024Updated 2 years ago
- Official Repo for LayerCraft☆19May 3, 2026Updated 4 months ago
- Code for paper "Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI"☆13Jan 19, 2024Updated 2 years ago
- This repository contains the source codes for the paper: "SPACE: A Simulator for Physical Interactions and Causal Learning in 3D Environm…☆16Oct 11, 2021Updated 4 years ago
- Train deepseek r1-like reasoning LLM with ease | 轻松训练1个deepseek r1类的推理LLM☆21Feb 15, 2025Updated last year
- Repository of paper: Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models☆36Sep 19, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆13Jun 13, 2023Updated 3 years ago
- [ICCV2025] ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors☆60May 17, 2026Updated 4 months ago
- Official codes of "Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs"☆19Feb 15, 2026Updated 7 months ago
- [EMNLP 2025] AutoSteer: Automating Steering for Safe Multimodal Large Language Models☆16Aug 21, 2025Updated last year
- Official resource for paper Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models (ACL 20…☆18Aug 12, 2024Updated 2 years ago
- Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models☆45Jun 14, 2024Updated 2 years ago
- [ICML 2026 Spotlight] UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models☆29Sep 9, 2026Updated last week