Official repo of From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
☆27Jun 23, 2026Updated 3 months ago
Alternatives and similar repositories for OSI-Bench
Users that are interested in OSI-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026 Spotlight] Code for miXed Discrete Diffusion Language Model☆31Mar 16, 2026Updated 6 months ago
- [AAAI2025] ChatterBox: Multi-round Multimodal Referring and Grounding, Multimodal, Multi-round dialogues☆62May 2, 2025Updated last year
- The official implementation of MaskGRPO: Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models. (ICLR 2026, arxiv…☆19Jan 27, 2026Updated 8 months ago
- official repo for `thinking with images through-self-calling`☆26Dec 28, 2025Updated 9 months ago
- Implementation of paper "CC-Diff: Enhancing Contextual Coherence in Remote Sensing Image Synthesis"☆29Dec 19, 2025Updated 9 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICLR 2026] Geometric-Mean Policy Optimization☆105Jan 26, 2026Updated 8 months ago
- [CVPR 2026] LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding☆57Jul 7, 2026Updated 3 months ago
- [ICCV 2023] Generative Prompt Model for Weakly Supervised Object Localization☆57Nov 10, 2023Updated 2 years ago
- Office implementation of "3D-Aware Hypothesis & Verification for Generalizable Relative Object Pose Estimation", ICLR 2024☆12Nov 5, 2024Updated last year
- CVPR2024, Semantic-aware SAM for Point-Prompted Instance Segmentation☆37Jan 20, 2025Updated last year
- [CVPR 2025] Adaptive Keyframe Sampling for Long Video Understanding☆238Dec 19, 2025Updated 9 months ago
- (CVPR2023/TPAMI2024) Integrally Pre-Trained Transformer Pyramid Networks -- A Hierarchical Vision Transformer for Masked Image Modeling☆218Jul 28, 2024Updated 2 years ago
- ☆76Mar 1, 2023Updated 3 years ago
- [ECCV 2024] ControlCap: Controllable Region-level Captioning☆81Oct 25, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [CVPR 2026] Official repository of "StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synth…☆17Feb 21, 2026Updated 7 months ago
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆20Jul 4, 2025Updated last year
- Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation.☆16Jul 21, 2025Updated last year
- Code for the paper "Data Attribution for Text-to-Image Models by Unlearning Synthesized Images."☆17May 23, 2025Updated last year
- Holistic Evaluation of Multimodal LLMs on Spatial Intelligence☆131Jul 1, 2026Updated 3 months ago
- Official PyTorch Implementation of "Better Source, Better Flow: Learning Condition-Dependent Source Distribution for Flow Matching"☆45Sep 24, 2026Updated 2 weeks ago
- The repository of the ACCV 2024 paper "FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Ge…☆13Aug 15, 2026Updated last month
- This is an official repository for the paper, NoiseCollage, which is a revolutionary extension of text-to-image diffusion models for layo…☆64May 16, 2024Updated 2 years ago
- Code repository for MMUGL: Multi-modal Graph Learning over UMLS Knowledge Graphs☆11Dec 7, 2023Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official Repo for LayerCraft☆19May 3, 2026Updated 5 months ago
- Code for paper "Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI"☆13Jan 19, 2024Updated 2 years ago
- This repository contains the source codes for the paper: "SPACE: A Simulator for Physical Interactions and Causal Learning in 3D Environm…☆15Oct 11, 2021Updated 5 years ago
- Train deepseek r1-like reasoning LLM with ease | 轻松训练1个deepseek r1类的推理LLM☆21Feb 15, 2025Updated last year
- A simple Python implement of Bilateral Mesh Denoising☆10Dec 1, 2019Updated 6 years ago
- This is an implementation of the paper "Are We Done with Object-Centric Learning?"☆14Jun 21, 2026Updated 3 months ago
- ☆13Jun 13, 2023Updated 3 years ago
- [ICCV2023] Spatio-temporal Prompting Network for Robust Video Feature Extraction☆11Aug 17, 2023Updated 3 years ago
- [EMNLP 2025] AutoSteer: Automating Steering for Safe Multimodal Large Language Models☆16Aug 21, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models☆45Jun 14, 2024Updated 2 years ago
- [ICML 2026 Spotlight] UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models☆30Sep 9, 2026Updated last month
- Automated detection of exudates from fundus images plays an important role in diabetic retinopathy (DR) screening and evaluation, for whi…☆11Dec 11, 2020Updated 5 years ago
- ☆12Apr 13, 2019Updated 7 years ago
- DASH: Detection and Assessment of Systematic Hallucinations of VLMs☆16Jul 2, 2025Updated last year
- THEORY OF SPACE: a benchmark for evaluating whether foundation models can actively explore under partial observability efficiently to bui…☆87Feb 27, 2026Updated 7 months ago
- 哈尔滨工业大学(深圳)2021年球季学期深度学习体系结构实验☆17Oct 1, 2022Updated 4 years ago