Collect the awesome works evolved around reasoning models like O1/R1 in visual domain
☆55Jul 21, 2025Updated last year
Alternatives and similar repositories for awesome-deep-multimodal-reasoning
Users that are interested in awesome-deep-multimodal-reasoning are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual in…☆1,502Mar 9, 2026Updated 6 months ago
- [NeurIPS 2025] NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation☆112Sep 18, 2025Updated 11 months ago
- Arbitrary Entropy Policy Optimization: Entropy Is Controllable in Reinforcement Fine-tuning☆18Jan 19, 2026Updated 7 months ago
- This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-bas…☆1,441Aug 2, 2026Updated last month
- ☆16Aug 19, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This is the repository for TNNLS paper: "Unihead: unifying multi-perception for detection heads"☆16Jan 13, 2025Updated last year
- ☆24Nov 4, 2025Updated 10 months ago
- Official implementation of Our NeurIPS 2024 Paper "Boundary Matters: A Bi-Level Active Finetuning Method"☆14Feb 11, 2025Updated last year
- Checkpoints, logs and source code for AAAI-23 paper 'Data-Efficient Image Quality Assessment with Attention-Panel Decoder'☆39Apr 3, 2024Updated 2 years ago
- Hypergraph Multi-Modal Learning for EEG-based Emotion Recognition in Conversation☆19Jun 15, 2026Updated 2 months ago
- [ACL'26 Oral] AgentOCR is a token-efficient framework that compresses multi-turn agent history by rendering it into images and adopting R…☆49Mar 1, 2026Updated 6 months ago
- [EMNLP25 Main]The official code of "Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval"☆26Mar 30, 2026Updated 5 months ago
- [ECCV2024] Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models☆20Jul 17, 2024Updated 2 years ago
- [ICML 2024] PyTorch implementation for "Diversified Batch Selection for Training Acceleration"☆10Jul 30, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆12Jul 31, 2024Updated 2 years ago
- [ICCV 2023] This is for the paper "Deep Homography Mixture for Single Image Rolling Shutter Correction".☆17May 25, 2025Updated last year
- The official codes for "Can Modern LLMs Act as Agent Cores in Radiology Environments?"☆32Aug 28, 2026Updated 2 weeks ago
- ☆11Jun 28, 2020Updated 6 years ago
- 关于LLM和Multimodal LLM的paper list☆64Aug 20, 2026Updated 3 weeks ago
- The notes for pytorch geometric learning☆10Jul 5, 2020Updated 6 years ago
- [NeurIPS 2023] Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language Models☆23Oct 21, 2025Updated 10 months ago
- Official implementation for "Pure Noise to the Rescue of Insufficient Data: Improving Imbalanced Classification by Training on Random Noi…☆15Jun 11, 2022Updated 4 years ago
- [Blog 1] Recording a bug of grpo_trainer in some R1 projects☆23Feb 23, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Accepted by CVPR 2022☆36May 23, 2022Updated 4 years ago
- A collection of awesome resources on image-to-image translation with diffusion models.☆16Mar 5, 2023Updated 3 years ago
- This is the official GitHub repository for our survey paper "Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language …☆208Jul 11, 2026Updated 2 months ago
- [TMLR 2025] Efficient Reasoning Models: A Survey☆320Jun 26, 2026Updated 2 months ago
- The official code of "Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search"☆38Jul 25, 2026Updated last month
- text-only training or language-free training for multimodal tasks (image/audio/video caption, retrieval, text2image)☆13Oct 15, 2024Updated last year
- [CVPR 2025] Official Pytorch Code for Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synth…☆15Jun 21, 2025Updated last year
- Official PyTorch Implementation for CAiD: Context-Aware Instance Discrimination for Self-supervised Learning in Medical Imaging - MIDL 20…☆11Apr 15, 2022Updated 4 years ago
- A curated list of researches in object-centric learning☆12Oct 14, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- OpenThinkIMG is an end-to-end open-source framework that empowers LVLMs to think with images.☆405Jun 1, 2025Updated last year
- ROSE: Robust Cross Supervision with Neighborhood Mining for Source-free Graph Domain Adaptation☆20Oct 22, 2024Updated last year
- [CVPR 2025] Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure Priors☆21Jun 6, 2025Updated last year
- The official code of "VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning" [NeurIPS25]☆192Jun 5, 2025Updated last year
- Agentic MLLMs☆215Oct 24, 2025Updated 10 months ago
- Code for "ConsistentChat: Building Skeleton-Guided Consistent Multi-Turn Dialogues for Large Language Models from Scratch", where dataset…☆17Sep 8, 2025Updated last year
- An implementation for Generator Versus Segmentor: Pseudo-healthy Synthesis☆12Oct 22, 2021Updated 4 years ago