[NeurIPS 2025] MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
☆108Sep 19, 2025Updated 11 months ago
Alternatives and similar repositories for MINT-CoT
Users that are interested in MINT-CoT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆14Mar 18, 2025Updated last year
- [CVPR' 25] Interleaved-Modal Chain-of-Thought☆112Dec 30, 2025Updated 8 months ago
- The first Interleaved framework for textual reasoning within the visual generation process☆167Mar 16, 2026Updated 5 months ago
- OpenCoF: Learning to Reason Through Video Generation☆76Jul 10, 2026Updated 2 months ago
- [CVPR 2026] Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens☆298Aug 2, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆18Apr 9, 2026Updated 5 months ago
- Official Code for "Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search"☆425Jan 29, 2026Updated 7 months ago
- MME-CoT: Benchmarking Chain-of-Thought in LMMs for Reasoning Quality, Robustness, and Efficiency☆135Aug 5, 2025Updated last year
- Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual in…☆1,502Mar 9, 2026Updated 6 months ago
- ☆73Feb 1, 2026Updated 7 months ago
- ☆1,272Nov 20, 2025Updated 9 months ago
- Are Video Models Ready as Zero-shot Reasoners?☆88Nov 24, 2025Updated 9 months ago
- Pixel-Level Reasoning Model trained with RL [NeuIPS25]☆307Jul 28, 2026Updated last month
- ☆17May 19, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official codebase for the paper Latent Visual Reasoning☆180Oct 22, 2025Updated 10 months ago
- [CVPR 2026] Official codes of "Monet: Reasoning in Latent Visual Space Beyond Image and Language"☆223Mar 19, 2026Updated 5 months ago
- [ICLR 2026] The official repository for paper "ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning"☆198May 1, 2026Updated 4 months ago
- ☆15Apr 20, 2026Updated 4 months ago
- [ICLR 2025] Mathematical Visual Instruction Tuning for Multi-modal Large Language Models☆157Dec 5, 2024Updated last year
- One Discrete Word for Visual Reasoning Overtakes Agentic and Latent Methods☆139Jun 9, 2026Updated 3 months ago
- [ICLR 2026]The official implementation of The paper "Exploring the Potential of Encoder-free Architectures in 3D LMMs"☆11Jan 26, 2026Updated 7 months ago
- [ECCV 2024] Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?☆183Apr 28, 2025Updated last year
- Offical Repository for Paper: DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation☆19Dec 7, 2025Updated 9 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms☆25Dec 21, 2025Updated 8 months ago
- [EMNLP Main 2026]VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning.☆26Jul 20, 2026Updated last month
- [NeurIPS'25] The official code of "PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning"☆30Mar 30, 2026Updated 5 months ago
- [EMNLP 2024 Findings] ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs☆30May 22, 2025Updated last year
- ☆35Feb 12, 2026Updated 6 months ago
- [ICML 2026] Official implementation of "Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence"☆163May 1, 2026Updated 4 months ago
- [ICML 2024] SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models☆22May 28, 2024Updated 2 years ago
- ☆19Jul 21, 2025Updated last year
- [ICLR'26] Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology☆93Jan 26, 2026Updated 7 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey☆1,025May 22, 2026Updated 3 months ago
- This repository is the official implementation of "Look-Back: Implicit Visual Re-focusing in MLLM Reasoning".☆98Jul 10, 2025Updated last year
- ☆137Jul 22, 2025Updated last year
- [MTI-LLM@NeurIPS 2025] Official implementation of "PyVision: Agentic Vision with Dynamic Tooling."☆163Jul 22, 2025Updated last year
- v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning☆21Updated this week
- The official code of "VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning" [NeurIPS25]☆192Jun 5, 2025Updated last year
- This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-bas…☆1,440Aug 2, 2026Updated last month