A curated list of the latest advancements, papers, tools, and datasets for **Multimodal Retrieval-Augmented Generation (RAG)**. Multimodal RAG integrates information retrieval and generation across multiple data modalities (e.g., text, image, video, audio).
☆53Sep 17, 2026Updated this week
Alternatives and similar repositories for Awesome-Multimodal-RAG
Users that are interested in Awesome-Multimodal-RAG are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [CVPR2025] VDocRAG: Retirval-Augmented Generation over Visually-Rich Documents☆68May 26, 2025Updated last year
- UniPrompt provides a unified interface to prompt optimization. We have distilled common functions from different algorithms and provide a…☆20May 20, 2025Updated last year
- Official Repo for paper "VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference"☆15Mar 28, 2026Updated 5 months ago
- CatRAG is a RAG framework builds on the HippoRAG 2 architecture and transforms the static KG into query-adaptive navigation structure. RA…☆29Aug 20, 2026Updated last month
- [ACL 2025 (Findings)] DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling☆22Dec 16, 2024Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Code and data for the EMNLP 2021 paper "Just Say No: Analyzing the Stance of Neural Dialogue Generation in Offensive Contexts". Coming so…☆17Jul 27, 2023Updated 3 years ago
- The code used to train and run inference with MMDocRAG☆22Nov 6, 2025Updated 10 months ago
- Official Implementation of LatentSwap3D: Semantic Edits on 3D Image GANs☆23Nov 28, 2023Updated 2 years ago
- Dataset and scripts for HRDoc☆42Jun 21, 2023Updated 3 years ago
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- Official PyTorch implementation of the TMI paper "Nucleus-aware Self-supervised Pretraining Using Unpaired Image-to-image Translation for…☆17Mar 13, 2024Updated 2 years ago
- Accelerating GOT-OCRv2 with VLLM☆10Nov 15, 2024Updated last year
- 助力搭建你的数字 AI 员工军团 — 多 Agent 一行命令组网协作。Claude Code / Claude Agent SDK / Codex / Grok Build 4 runtime + 8+ 家 LLM(Anthropic / OpenAI / xAI / M…☆74Updated this week
- ☆26Aug 24, 2022Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆14May 16, 2025Updated last year
- Official code for the ICLR 2025 paper, "Ada-K Routing: Boosting the Efficiency of MoE-based LLMs"☆12Mar 1, 2025Updated last year
- ☆17Mar 9, 2026Updated 6 months ago
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents☆45Apr 13, 2026Updated 5 months ago
- ☆10Sep 17, 2025Updated last year
- MemSyco-Bench: Benchmarking Sycophancy in Agent Memory☆20Sep 11, 2026Updated last week
- Implementation and evaluation of multimodal RAG with text and image inputs for industrial applications☆72Nov 6, 2024Updated last year
- Highly Efficient Query Rewriter for Passage Retrieval in the realm of Retrieval-Augmented Generation (RAG)☆31Aug 15, 2026Updated last month
- 在RAG技术中,嵌入向量的生成和匹配是关键环节。本文介绍了一种基于CLIP/BLIP模型的嵌入服务,该服务支持文本和图像的嵌入生成与相似度计算,为多模态信息检索提供了基础能力。☆42Dec 28, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ACM MM2025] Official code of " HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation"☆113Jul 23, 2025Updated last year
- Evaluation benchmark for the task of Semantic Image Translation. Contains code to run FlexIT (CVPR 2022)☆34Mar 25, 2022Updated 4 years ago
- Utility functions/scripts for working with GPUs.☆10Jul 5, 2021Updated 5 years ago
- [ACL 2025] Towards Text-Image Interleaved Retrieval☆16Sep 3, 2025Updated last year
- [ICLR 2020] Haotao Wang, Tianlong Chen, Zhangyang Wang, Kede Ma, "I Am Going MAD: Maximum Discrepancy Competition for Comparing Classifie…☆20Dec 30, 2021Updated 4 years ago
- [ACL 2024 Oral] This is the code repo for our ACL‘24 paper "MARVEL: Unlocking the Multi-Modal Capability of Dense Retrieval via Visual Mo…☆39Jun 30, 2024Updated 2 years ago
- Multimodal Retrieval-augmented Generation Framework Built by Tongyi Lab, Alibaba Group.☆985Apr 29, 2026Updated 4 months ago
- [ICDAR 2024] (Best Student Paper🏆) Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation☆14Sep 6, 2024Updated 2 years ago
- Learning Ontologies Via Embeddings☆12Jul 6, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Awesome-RAG-Vision: a curated list of advanced retrieval augmented generation (RAG) for Computer Vision☆341Jan 25, 2026Updated 7 months ago
- This is the official repository for Retrieval Augmented Visual Question Answering☆253Dec 19, 2024Updated last year
- Contrastive self-supervised learning using Rényi divergence☆14Oct 21, 2022Updated 3 years ago
- LONGAGENT: Scaling Language Models to 128k Context through Multi-Agent Collaboration☆11Mar 11, 2024Updated 2 years ago
- To mitigate position bias in LLMs, especially in long-context scenarios, we scale only one dimension of LLMs, reducing position bias and …☆12Jun 18, 2024Updated 2 years ago
- Official repository for the paper "Reconstruction of Perceived Images from fMRI Patterns and Semantic Brain Exploration using Instance-Co…☆24May 19, 2022Updated 4 years ago
- ☆12Jan 10, 2025Updated last year