Collect the awesome works evolved around reasoning models like O1/R1 in visual domain
☆55Jul 21, 2025Updated last year
Alternatives and similar repositories for awesome-deep-multimodal-reasoning
Users that are interested in awesome-deep-multimodal-reasoning are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual in…☆1,502Mar 9, 2026Updated 5 months ago
- EPIC-Bench is a fine-grained embodied visual grounding benchmark for evaluating VLMs on target localization, navigation-oriented percepti…☆16May 28, 2026Updated 2 months ago
- [NeurIPS 2025] NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation☆112Sep 18, 2025Updated 10 months ago
- [ACM MM2026] This is the official implementation of MedCCO☆17Jul 12, 2026Updated last month
- This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-bas…☆1,439Aug 2, 2026Updated last week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Official repo for "PAPO: Perception-Aware Policy Optimization for Multimodal Reasoning"☆156Feb 4, 2026Updated 6 months ago
- ☆16Sep 16, 2025Updated 10 months ago
- This is the repository for TNNLS paper: "Unihead: unifying multi-perception for detection heads"☆16Jan 13, 2025Updated last year
- ☆24Nov 4, 2025Updated 9 months ago
- DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use☆28Mar 13, 2026Updated 4 months ago
- [ECCV 2024] Soft Prompt Generation for Domain Generalization☆33Oct 1, 2024Updated last year
- [ACL'26 Oral] AgentOCR is a token-efficient framework that compresses multi-turn agent history by rendering it into images and adopting R…☆43Mar 1, 2026Updated 5 months ago
- [EMNLP25 Main]The official code of "Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval"☆25Mar 30, 2026Updated 4 months ago
- ☆17May 14, 2026Updated 2 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- [ICLR 2026]🚀ReVisual-R1 is a 7B open-source multimodal language model that follows a three-stage curriculum—cold-start pre-training, mul…☆211Dec 10, 2025Updated 8 months ago
- [ECCV2024] Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models☆20Jul 17, 2024Updated 2 years ago
- EgoToM is an egocentric theory-of-mind benchmark built on Ego4D videos, containing multi-choice questions that evaluate multimodal large …☆16Apr 1, 2025Updated last year
- [ICML 2026] Milestone-Guided Policy Learning for Long-Horizon Language Agents☆42May 29, 2026Updated 2 months ago
- The official codes for "Can Modern LLMs Act as Agent Cores in Radiology Environments?"☆29Jan 22, 2025Updated last year
- [ICCV 2023] This is for the paper "Deep Homography Mixture for Single Image Rolling Shutter Correction".☆17May 25, 2025Updated last year
- Albert for Conversational Question Answering Challenge☆21Jun 12, 2023Updated 3 years ago
- ☆11Oct 2, 2024Updated last year
- Code for "LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language Model", CVPR 2024 Highlight☆65Jun 11, 2024Updated 2 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- KARL: Knowledge-Aware Reasoning and Reinforcement Learning for Knowledge-Intensive Visual Grounding☆70Apr 5, 2026Updated 4 months ago
- 关于LLM和Multimodal LLM的paper list☆64Jun 17, 2026Updated last month
- [NeurIPS 2023] Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language Models☆23Oct 21, 2025Updated 9 months ago
- Official implementation for "Pure Noise to the Rescue of Insufficient Data: Improving Imbalanced Classification by Training on Random Noi…☆15Jun 11, 2022Updated 4 years ago
- [Blog 1] Recording a bug of grpo_trainer in some R1 projects☆23Feb 23, 2025Updated last year
- Code of IEEE TIM Paper: ETDNet: Efficient Transformer-Based Detection Network for Surface Defect Detection☆29Oct 16, 2023Updated 2 years ago
- ☆25Aug 1, 2025Updated last year
- [TMLR 2025] Efficient Reasoning Models: A Survey☆318Jun 26, 2026Updated last month
- The official code of "Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search"☆36Jul 25, 2026Updated 2 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- text-only training or language-free training for multimodal tasks (image/audio/video caption, retrieval, text2image)☆13Oct 15, 2024Updated last year
- Official PyTorch Implementation for CAiD: Context-Aware Instance Discrimination for Self-supervised Learning in Medical Imaging - MIDL 20…☆11Apr 15, 2022Updated 4 years ago
- [ACM MM 2024] Improving Composed Image Retrieval via Contrastive Learning with Scaling Positives and Negatives☆39Sep 9, 2025Updated 11 months ago
- Edge-oriented Point cloud Transformer for 3D Intracranial Aneurysm Segmentation. MICCAI22☆13Aug 18, 2022Updated 3 years ago
- A curated list of researches in object-centric learning☆12Oct 14, 2024Updated last year
- OpenThinkIMG is an end-to-end open-source framework that empowers LVLMs to think with images.☆400Jun 1, 2025Updated last year
- [CVPR 2025] Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure Priors☆21Jun 6, 2025Updated last year