[ICML2026] Imagination Helps Visual Reasoning, But Not Yet in Latent Space
☆29May 4, 2026Updated 3 months ago
Alternatives and similar repositories for CapImagine
Users that are interested in CapImagine are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [EMNLP'25 main] Official Implementation of ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guid…☆27Feb 8, 2026Updated 5 months ago
- [ICML 2026] Official implementation of Vision-aligned Latent Reasoning for Multi-modal Large Language Model (VaLR)☆20Apr 30, 2026Updated 3 months ago
- [CVPR2026] Official codebase for the paper "Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space"☆85May 12, 2026Updated 2 months ago
- Official codebase for the paper Latent Visual Reasoning☆172Oct 22, 2025Updated 9 months ago
- This repository is the official implementation of "DG-Mamba: Robust and Efficient Dynamic Graph Structure Learning with Selective State S…☆23Apr 17, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Companion repository to "Prompt Compression and Contrastive Conditioning for Controllability and Toxicity Reduction in Language Models"☆14May 31, 2023Updated 3 years ago
- [ACL 2026] Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning☆93Jan 22, 2026Updated 6 months ago
- This repository contains a regularly updated paper list for LLMs-reasoning-in-latent-space.☆369Jun 20, 2026Updated last month
- [CVPR 2026] Official codes of "Monet: Reasoning in Latent Visual Space Beyond Image and Language"☆216Mar 19, 2026Updated 4 months ago
- [NeurIPS 2025] Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains☆98Jun 29, 2026Updated last month
- [ICLR 2025] Official Implementation of Local-Prompt: Extensible Local Prompts for Few-Shot Out-of-Distribution Detection☆52Jul 30, 2025Updated last year
- [ICLR2026] Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception☆17Jan 26, 2026Updated 6 months ago
- [NeurIPS 2025] Official code for paper: Latent Chain-of-Thought for Visual Reasoning☆36Oct 16, 2025Updated 9 months ago
- Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents☆31Apr 16, 2026Updated 3 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- YAICON 3rd project page - 4D Gaussian for Head Reconstruction☆11Dec 22, 2023Updated 2 years ago
- ☆22Jul 26, 2026Updated last week
- [ECCV 2026] Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models☆28Jun 20, 2026Updated last month
- [ICLR 2026] Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization☆32Mar 6, 2026Updated 4 months ago
- ☆24May 23, 2025Updated last year
- [CVPR 2026] Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens☆294Aug 2, 2025Updated last year
- KARL: Knowledge-Aware Reasoning and Reinforcement Learning for Knowledge-Intensive Visual Grounding☆69Apr 5, 2026Updated 4 months ago
- Official code for our paper "Model Composition for Multimodal Large Language Models" (ACL 2024)☆31Jan 8, 2025Updated last year
- [ICLR'26] "Nabla-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space" by Peihao Wang*, Ruisi Cai*, Zhen Wang, Hongyuan…☆35Mar 10, 2026Updated 4 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- 中国科学院大学研究生课程-高级人工智能☆10Jan 8, 2022Updated 4 years ago
- [ICLR26] Understanding VS. Generation: Navigating Optimization Dilemma in Multimodal Models☆26May 6, 2026Updated 2 months ago
- [ICCV 2025] Official Implementation of Federated Continual Instruction Tuning☆17Aug 10, 2025Updated 11 months ago
- [CVPR2026] CodePercept: Code-Grounded Visual STEM Perception for MLLM☆44Jul 28, 2026Updated last week
- [ACL2025 Findings] Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models☆91May 20, 2025Updated last year
- [NeurIPS 2024] "Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection"☆13Oct 28, 2024Updated last year
- Prior and Prediction Inverse Kernel Transformer for Single Image Defocus Deblurring☆11Mar 12, 2024Updated 2 years ago
- Build coherent and visually polished multimodal webpages with hierarchical planning, AIGC tools, and iterative reflection.☆16May 17, 2026Updated 2 months ago
- UniDoc-RL: Unified Document Understanding with Reinforcement Learning☆16May 21, 2026Updated 2 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Search Engine Guided Non-Parametric Neural Machine Translation☆14Oct 23, 2017Updated 8 years ago
- ☆15Jan 12, 2026Updated 6 months ago
- Repo for paper "CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models".☆13Oct 14, 2024Updated last year
- ☆25Apr 7, 2025Updated last year
- ☆16Jun 17, 2026Updated last month
- Multilingual Neural Machine Translation using Transformers with Conditional Normalization.☆18Mar 24, 2023Updated 3 years ago
- ☆14Jul 17, 2024Updated 2 years ago