Latest open-source "Thinking with images" (O3/O4-mini) papers, covering training-free, SFT-based, and RL-enhanced methods for "fine-grained visual understanding".
☆115Aug 21, 2025Updated 11 months ago
Alternatives and similar repositories for Awesome-Thinking-With-Images
Users that are interested in Awesome-Thinking-With-Images are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual in…☆1,498Mar 9, 2026Updated 4 months ago
- Interleaving Reasoning: Next-Generation Reasoning Systems for AGI☆280Jun 5, 2026Updated last month
- ☆1,255Nov 20, 2025Updated 8 months ago
- This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-bas…☆1,438May 11, 2026Updated 2 months ago
- [ICCV 2023] Simple Baselines for Interactive Video Retrieval with Questions and Answers☆20Apr 16, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning☆36Jun 10, 2025Updated last year
- Latest Advances on (RL based) Multimodal Reasoning and Generation in Multimodal LLMs☆83Updated this week
- Reading notes about Multimodal Large Language Models, Large Language Models, and Diffusion Models☆1,187Jul 14, 2026Updated 2 weeks ago
- Official Code for "Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search"☆423Jan 29, 2026Updated 6 months ago
- Pixel-Level Reasoning Model trained with RL [NeuIPS25]☆301Jun 4, 2026Updated last month
- ☆19Sep 19, 2024Updated last year
- [AAAI 2025] Open-vocabulary Video Instance Segmentation Codebase built upon Detectron2, which is really easy to use.☆26Dec 30, 2024Updated last year
- ☆12Mar 22, 2025Updated last year
- Official PyTorch implementation of RACRO (https://www.arxiv.org/abs/2506.04559)☆19Jul 1, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Official eval code for ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation☆27Dec 12, 2025Updated 7 months ago
- EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL☆5,092Updated this week
- [NeurIPS 2025] Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing☆97Jul 27, 2025Updated last year
- Repository for awesome spatial/visual reasoning MLLMs. (focus more on embodied applications)☆71Jun 26, 2025Updated last year
- The official repo for "Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation", ECCV 2024☆18Oct 11, 2024Updated last year
- [ACL2026 oral] Uni-MMMU : A Massive Multi-discipline Multimodal Unified Benchmark☆26Apr 13, 2026Updated 3 months ago
- ✨First Open-Source R1-like Video-LLM [2025/02/18]☆382Jul 1, 2026Updated 3 weeks ago
- ☆117Aug 14, 2025Updated 11 months ago
- ☆20Apr 21, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ICLR 2026] The official repo of "MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs"☆46Jul 3, 2026Updated 3 weeks ago
- ☆46Jul 14, 2025Updated last year
- [ Arxiv 2023 ] This repository contains the code for "MUPPET: Multi-Modal Few-Shot Temporal Action Detection"☆16Aug 30, 2023Updated 2 years ago
- ☆133Jul 22, 2025Updated last year
- RLHF for Video Diffusion Models☆26Jul 30, 2025Updated last year
- VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation [TMLR26]☆15Jun 1, 2026Updated 2 months ago
- Extend OpenRLHF to support LMM RL training for reproduction of DeepSeek-R1 on multimodal tasks.☆848May 14, 2025Updated last year
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey☆1,016May 22, 2026Updated 2 months ago
- [ICCV2025]Code Release of Harmonizing Visual Representations for Unified Multimodal Understanding and Generation☆192May 21, 2025Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Official repository of 'Visual-RFT: Visual Reinforcement Fine-Tuning' & 'Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning'’☆2,263Oct 29, 2025Updated 9 months ago
- ☆34Dec 18, 2025Updated 7 months ago
- [ACM MM 2025] ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models☆18Jul 15, 2025Updated last year
- [NeurIPS 2025] MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning☆107Sep 19, 2025Updated 10 months ago
- https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT☆137Jan 30, 2026Updated 6 months ago
- [TMLR 2025🔥] A survey for the autoregressive models in vision.☆805May 5, 2026Updated 2 months ago
- 📚 A curated collection of papers and open-source code repositories dedicated to the application of Vision-Language Models (VLMs) for str…☆190Jul 22, 2026Updated last week