A Survey of Multimodal Retrieval-Augmented Generation
☆20Nov 3, 2025Updated 10 months ago
Alternatives and similar repositories for MRAGSurvey
Users that are interested in MRAGSurvey are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This repository is the implementation of "QT-TextSR: Enhancing scene text image super-resolution via efficient interaction with text reco…☆21Jul 9, 2025Updated last year
- [MM'2024] Official release of RFUND introduced in the MM'2024 paper "PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking f…☆21Dec 4, 2024Updated last year
- -☆24Oct 25, 2022Updated 3 years ago
- SCUT-EnsExam is a real-world handwritten text erasure dataset for examination paper scenarios, which consists of 545 examination paper im…☆24Jul 17, 2026Updated 2 months ago
- ☆19Sep 11, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [ACM MM 2022] Marior: Margin Removal and Iterative Content Rectification for Document Dewarping in the Wild☆27Aug 12, 2022Updated 4 years ago
- ☆36Aug 1, 2026Updated last month
- Official implementation of URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding (AAAI 2026…☆43Feb 4, 2026Updated 7 months ago
- ☆71Aug 1, 2026Updated last month
- [EMNLP 2024] SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information☆11Oct 11, 2024Updated last year
- Official implementation of ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining (AAAI 20…☆70Jul 4, 2024Updated 2 years ago
- The dataset used in the CVPR 2022 paper (SimAN: Exploring Self-Supervised Representation Learning of Scene Text via Similarity-Aware Norm…☆34Jun 21, 2022Updated 4 years ago
- Official implementation of UPOCR: Towards unified pixel-level OCR interface (ICML 2024)☆73Jun 6, 2024Updated 2 years ago
- On the Robustness of GUI Grounding Models Against Image Attacks☆12Apr 8, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICLR 2026] OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning☆78May 26, 2026Updated 3 months ago
- Implementation and evaluation of multimodal RAG with text and image inputs for industrial applications☆72Nov 6, 2024Updated last year
- [CVPR 2025] DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding☆29Dec 18, 2025Updated 9 months ago
- An unofficial PyTorch implementation of "Lin et al. ViBERTgrid: A Jointly Trained Multi-Modal 2D Document Representation for Key Informat…☆53Jan 9, 2024Updated 2 years ago
- The official repo for our paper: LegalAgentBench: Evaluating LLM Agents in Legal Domainl☆50Apr 10, 2026Updated 5 months ago
- Code and data for the paper: DTSM: Toward Dense Table Structure Recognition with Text Query Encoder and Adjacent Feature Aggregator☆14Apr 28, 2024Updated 2 years ago
- Confidence Regulation Neurons in Language Models (NeurIPS 2024)☆16Feb 1, 2025Updated last year
- [DASFAA 2025] Enhancing Retrieval-Augmented Generation with Multi-Modal Knowledge Graph Integration☆15Feb 28, 2026Updated 6 months ago
- ☆53May 11, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [IJCAI'21] Code for Context-Aware Image Inpainting with Learned Semantic Priors,☆57Dec 9, 2021Updated 4 years ago
- ☆19Dec 10, 2023Updated 2 years ago
- A toolbox for EEG signals processing. Welcome to join and build!☆13Nov 9, 2022Updated 3 years ago
- (ICCV 2023) ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in Transformer☆78Apr 9, 2024Updated 2 years ago
- Official PyTorch implementation of the CVPR 2022 paper: "Look Closer to Supervise Better: One-Shot Font Generation via Component-Based Di…☆94Sep 17, 2022Updated 4 years ago
- The official implement of CTRNet++.☆15Dec 30, 2024Updated last year
- 该项目为对llama2进行微调及使用中文微调的技术细节,适合初学者观看。☆23Nov 9, 2023Updated 2 years ago
- The tampered text detection dataset☆22Aug 23, 2023Updated 3 years ago
- ☆17Sep 22, 2023Updated 3 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- 强化学习课程,主要是如何用强化学习解决问题☆15Dec 10, 2024Updated last year
- 本项目是关于Yi的多模态系列模型,如Yi-VL-6B/34B等的实验与应用。☆14Jan 25, 2024Updated 2 years ago
- ☆13Jul 16, 2024Updated 2 years ago
- [MM'2024] PEneo, an effective algorithm for key-value pair extraction from form-like documents, designed for real-world applications.☆41Apr 7, 2025Updated last year
- Code for ICCV2025 paper——IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves☆18Jul 11, 2025Updated last year
- official impelmentation of Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input☆67Aug 30, 2024Updated 2 years ago
- ABC: Achieving Better Control of Multimodal Embeddings using VLMs [TMLR2025]☆20Aug 21, 2025Updated last year