[AAAI2026] Source code for RegionRAG
☆26Apr 20, 2026Updated 5 months ago
Alternatives and similar repositories for RegionRAG
Users that are interested in RegionRAG are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [CVPR2025] Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models☆21Apr 30, 2025Updated last year
- What Is a Good Caption? A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness☆28May 16, 2025Updated last year
- The code used to train and run inference with the ColQwen3 model. Welcome to follow and star! ⭐️⭐️⭐️ https://huggingface.co/goodman2001/…☆16Aug 16, 2026Updated last month
- [CVPR2025] VDocRAG: Retirval-Augmented Generation over Visually-Rich Documents☆67May 26, 2025Updated last year
- [EMNLP 2025] Official implementation for paper "MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrieval"☆27Mar 17, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆12Nov 26, 2024Updated last year
- ☆37Apr 1, 2026Updated 6 months ago
- [IJCAI-2024] The official code of Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition☆10Aug 10, 2025Updated last year
- Jina VDR is a multilingual, multi-domain benchmark for visual document retrieval☆38Aug 4, 2025Updated last year
- Embedding model prioritized towards Multimodal RAG, overall + VisDoc double top1 on MMEB benchmark☆37Jun 16, 2026Updated 3 months ago
- Software Engineering Back End Microservices Project☆15Nov 20, 2024Updated last year
- ☆12Mar 31, 2025Updated last year
- [AAAI-26] Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?☆32Dec 14, 2025Updated 9 months ago
- MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding☆357Aug 8, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning☆20May 27, 2026Updated 4 months ago
- StelLA: Subspace Learning in Low-rank Adaptation using Stiefel Manifold (NeurIPS 2025 Spotlight)☆21Jun 29, 2026Updated 3 months ago
- The official code of "CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval"☆15Sep 19, 2024Updated 2 years ago
- Official PyTorch Implementation of ZSLViT (CVPR'24)☆18Jul 9, 2024Updated 2 years ago
- Phát triển ứng dụng web☆14Jan 7, 2022Updated 4 years ago
- SURDS: Self-Supervised Attention-guided Reconstruction and Dual Triplet Loss for Writer Independent Offline Signature Verification", ICPR…☆17Jul 22, 2022Updated 4 years ago
- ☆15Oct 6, 2024Updated 2 years ago
- SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models☆15Feb 20, 2025Updated last year
- Implementation of MLLM-based Self-Vision-RAG models☆15Nov 30, 2025Updated 10 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- MMhops-R1: Multimodal Multi-hop Reasoning☆16Aug 17, 2026Updated last month
- [ACM MM 2025 🔥🔥 ] MIRA: A first-of-its-kind medical RAG framework that fuses image features and retrieved knowledge with dynamic contex…☆23Aug 28, 2025Updated last year
- AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation (EMNLP 2024 Findings)☆19Dec 30, 2024Updated last year
- ☆17Dec 25, 2023Updated 2 years ago
- Official PyTorch Implementation of TransZero (AAAI'22)☆87Jan 20, 2024Updated 2 years ago
- [CVPR 2025] Official Repository of the paper "On the Consistency of Video Large Language Models in Temporal Comprehension"☆16Oct 13, 2025Updated 11 months ago
- [NeurIPS 2025] AutoPrune, a general pruning method for LLM/VLM/VLA☆20Oct 7, 2025Updated last year
- The official code of Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment Retrieval (AAAI2024)☆32Mar 29, 2024Updated 2 years ago
- Parsing-free RAG supported by VLMs☆980Aug 9, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [NeurIPS 2025] SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models.☆18May 26, 2026Updated 4 months ago
- A composed retrieval project☆18Apr 9, 2026Updated 6 months ago
- ☆70May 19, 2025Updated last year
- ☆20Oct 13, 2025Updated 11 months ago
- Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark containing 100K image-text pairs for robust image-text …☆16Nov 20, 2025Updated 10 months ago
- A hierarchical multi-agent framework for exhaustive cross-document question answering.☆21Mar 14, 2026Updated 6 months ago
- MomentDiff: Generative Video Moment Retrieval from Random to Real--NeurIPS 2023☆80Nov 2, 2023Updated 2 years ago