[CVPR 2026] Boosting Reasoning in Large Multimodal Models via Activation Replay
☆26Aug 25, 2026Updated last month
Alternatives and similar repositories for replay
Users that are interested in replay are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TokenAR: Multiple Subject Generation via Autoregressive Token-level enhancement☆22Aug 4, 2026Updated last month
- Official code repository for Med-CMR : "A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multi…☆28Dec 10, 2025Updated 9 months ago
- [ICML 2026] The official code of FeRA: Frequency–Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning☆29Dec 27, 2025Updated 9 months ago
- From Large Angles to Consistent Faces: Identity-Preserving Video Generation via Mixture of Facial Experts☆27Jan 12, 2026Updated 8 months ago
- Official repository for "Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models"☆24Dec 2, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆18Jul 31, 2025Updated last year
- ☆23May 26, 2025Updated last year
- ☆29Nov 28, 2025Updated 10 months ago
- [ICLR 26] Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow☆47Oct 3, 2025Updated 11 months ago
- ☆96Feb 5, 2026Updated 7 months ago
- [ICML 2026] Transform Trained Transformer for Accelerating Native 4K Video Generation☆41Dec 16, 2025Updated 9 months ago
- [ICLR 26] Context Tokens are Anchors: Understanding the Repeat Curse in dMLLMs from an Information Flow Perspective☆34Mar 6, 2026Updated 6 months ago
- [CVPR 2026] Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation☆64Dec 16, 2025Updated 9 months ago
- CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal☆26May 25, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆32Jan 11, 2026Updated 8 months ago
- [EMNLP 2026 Findings] The official code of Refinement Provenance Inference: Detecting LLM-Refined Training Prompts from Model Behavior☆22Jan 6, 2026Updated 8 months ago
- ☆42Nov 12, 2025Updated 10 months ago
- Virtual Try-on based on the powerful Flux model☆27Dec 4, 2024Updated last year
- [ICRA 26] C^2ROPE: Causal Continuous Rotary Positional Encoding for 3D Large Multimodal-Models Reasoning☆28Feb 13, 2026Updated 7 months ago
- [CVPR 2025] Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention☆70Jul 16, 2024Updated 2 years ago
- ✨✨The Curse of Multi-Modalities (CMM): Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio☆55Jul 11, 2025Updated last year
- [ICLR 2026] AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size☆17Jan 28, 2026Updated 8 months ago
- [CVPR2025] Unveil Inversion and Invariance in Flow Transformer for Versatile Image Editing☆25Aug 23, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆15Apr 6, 2026Updated 5 months ago
- [NeurIPS 2023] and [ICLR 2024] for robustness certification.☆10Nov 30, 2024Updated last year
- [ICLR 2026 Oral] 🎉Hallucination Begins Where Saliency Drops☆69Feb 12, 2026Updated 7 months ago
- ☆16Oct 12, 2024Updated last year
- Official repository for the paper "PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset"☆35Jul 10, 2026Updated 2 months ago
- MedEyes: Learning Dynamic Focus for Diagnostic Evolution Like a Clinician☆19Mar 13, 2026Updated 6 months ago
- ☆31Nov 17, 2024Updated last year
- [CVPR2026 Highlight] FlexMem: Scaling the Long Video Understanding of MLLMs via Visual Memory Mechanism☆42Apr 10, 2026Updated 5 months ago
- [CVPR'26] VisPlay: Self-Evolving Vision-Language Models☆77Feb 25, 2026Updated 7 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Code for ICLR 2025 Paper: Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs☆25May 7, 2025Updated last year
- Memory-optimized training scripts for video models based on Diffusers☆17Jan 3, 2025Updated last year
- A code base for the third place solution of Ego-Exo4D bodypose challenge for CVPR2024 workshop☆12Jun 16, 2024Updated 2 years ago
- [ICLR 2026] Official implementation of the paper "📷 On the Generalization Capacities of MLLMs for Spatial Intelligence"☆31Mar 17, 2026Updated 6 months ago
- Code Implementation for AutoAttend: Automated Attention Representation Search☆11Jul 26, 2021Updated 5 years ago
- [NeurIPS 2024] Mitigating Object Hallucination via Concentric Causal Attention☆69Aug 30, 2025Updated last year
- [CVPR 2024] The official implementation of paper "synthesize, diagnose, and optimize: towards fine-grained vision-language understanding"☆52Jun 16, 2025Updated last year