MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding
☆358Aug 8, 2025Updated last year
Alternatives and similar repositories for MDocAgent
Users that are interested in MDocAgent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [CVPR2025] VDocRAG: Retirval-Augmented Generation over Visually-Rich Documents☆67May 26, 2025Updated last year
- ☆16Jan 18, 2026Updated 8 months ago
- ☆70May 19, 2025Updated last year
- ☆46Jul 28, 2025Updated last year
- [ICLR'26] EduVisAgent: A Benchmark and Multi-Agent Framework for Pedagogical Visualization☆31Aug 5, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆42Jan 9, 2026Updated 8 months ago
- Parsing-free RAG supported by VLMs☆981Aug 9, 2026Updated last month
- ☆37Apr 1, 2026Updated 6 months ago
- ☆44Apr 6, 2026Updated 5 months ago
- [ICLR'25 Oral] MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models☆35Nov 3, 2024Updated last year
- Official Repository of MMLONGBENCH-DOC: Benchmarking Long-context Document Understanding with Visualizations☆157Sep 28, 2025Updated last year
- [ICML'25] MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization☆75Jun 5, 2025Updated last year
- [EMNLP 2025] ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents☆671Jan 11, 2026Updated 8 months ago
- [ACM MM2025] Official code of " HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation"☆113Jul 23, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official PyTorch Implementation of MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced …☆90Nov 15, 2024Updated last year
- A Survey on Multimodal Retrieval-Augmented Generation☆545Feb 20, 2026Updated 7 months ago
- ☆69Jul 1, 2026Updated 3 months ago
- [ICLR 2026] High-Fidelity Visual Reasoning on Structured Images☆36Jul 17, 2026Updated 2 months ago
- [ACL '25] Source code for our paper ''RankCoT: Refining Knowledge for Retrieval-Augmented Generation through Ranking Chain-of-Thoughts''☆54Nov 27, 2025Updated 10 months ago
- Multimodal Retrieval-augmented Generation Framework Built by Tongyi Lab, Alibaba Group.☆984Apr 29, 2026Updated 5 months ago
- UniDoc-RL: Unified Document Understanding with Reinforcement Learning☆18May 21, 2026Updated 4 months ago
- Official implementation of "Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation" (CVPR 202…☆39May 26, 2025Updated last year
- An implementation of "M3DOCRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding" by Jaemin Cho, Debanj…☆55Nov 13, 2024Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Implementation of MLLM-based Self-Vision-RAG models☆15Nov 30, 2025Updated 10 months ago
- ConfAgents: A Conformal-Guided Multi-Agent Framework for Cost-Efficient Medical Diagnosis☆15Jul 22, 2026Updated 2 months ago
- [ICML'26 & COLM'26] Agent0 Series: Self-Evolving Agents from Zero Data☆1,269Jul 10, 2026Updated 2 months ago
- [NeurIPS'26] SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning☆992Sep 25, 2026Updated last week
- Official repository for RAG-Gym☆128Jul 14, 2026Updated 2 months ago
- ☆22Dec 18, 2025Updated 9 months ago
- ✨✨[NeurIPS 2025] This is the official implementation of our paper "Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehensi…☆459Jun 26, 2026Updated 3 months ago
- Evaluation framework for paper "VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?"☆68Oct 19, 2024Updated last year
- ☆59Feb 27, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- The official repository of MM-R5☆29Jun 22, 2025Updated last year
- [EMNLP 2025] Official codebase for Rearank: Reasoning Re-ranking Agent☆39Aug 20, 2025Updated last year
- FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain (EMNLP 2025)☆22Jan 13, 2026Updated 8 months ago
- 🔍 Search-o1: Agentic Search-Enhanced Large Reasoning Models [EMNLP 2025]☆1,250Nov 17, 2025Updated 10 months ago
- [EMNLP26] Welcome! 😊 This is the official code release of EviNote-RAG, and we’re happy to share it with the community.☆48Aug 24, 2026Updated last month
- A universal skill ecosystem for AI agents.☆25Feb 6, 2026Updated 7 months ago
- The code used to train and run inference with the ColVision models, e.g. ColPali, ColQwen2, and ColSmol.☆2,823Sep 21, 2026Updated last week