[CVPR 2026 Highlight] DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
☆19Jun 4, 2026Updated 2 months ago
Alternatives and similar repositories for DocSeeker
Users that are interested in DocSeeker are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆27Jul 5, 2026Updated last month
- [Innovation 2026] Oracle bone script decipherment via human-workflow-inspired deep learning☆34Jul 22, 2026Updated last month
- ☆11Oct 31, 2024Updated last year
- [AAAI 2025] DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming☆36Jun 1, 2025Updated last year
- [ICLR 2026] OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning☆78May 26, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ACL '26] Lang2Act: Fine-Grained Visual Reasoning through Self-Emergent Linguistic Toolchains☆25Apr 7, 2026Updated 4 months ago
- Adapt MLLMs to Domains via Post-Training (EMNLP 2025 Findings)☆14Nov 11, 2025Updated 9 months ago
- UniDoc-RL: Unified Document Understanding with Reinforcement Learning☆17May 21, 2026Updated 3 months ago
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining☆55Apr 22, 2026Updated 4 months ago
- ☆14Jul 13, 2024Updated 2 years ago
- PureDocBench: source-traceable benchmark for document parsing across clean, degraded, and real-world settings☆41Jul 15, 2026Updated last month
- Unofficial Visual Prompt Tuning implementation☆17May 22, 2023Updated 3 years ago
- self ensemble label correction☆17Jul 29, 2022Updated 4 years ago
- Official code of the paper "Synthetic Instance Segmentation from Semantic Image Segmentation Masks"☆21Oct 31, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆12Mar 28, 2024Updated 2 years ago
- Official implementation for "Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation Decoders"☆25Aug 19, 2026Updated last week
- Source code of our TCSVT'22 paper Reading-strategy Inspired Visual Representation Learning for Text-to-Video Retrieval☆19Feb 13, 2022Updated 4 years ago
- The official code of "Towards Long-horizon Agentic Multimodal Search"☆29Apr 17, 2026Updated 4 months ago
- Graph-based experience memory for LLM reward prediction with limited labels. 20% labels → 97.3% Oracle.☆19Mar 24, 2026Updated 5 months ago
- Code for "Agentic Very Long Video Understanding" (EGAgent) [ACL 2026 Main]☆56Jul 1, 2026Updated last month
- Official evaluation scripts and baseline prompts for the DocVQA 2026 (ICDAR 2026) Competition on Multimodal Reasoning over Documents.☆18Mar 16, 2026Updated 5 months ago
- A python implement for Certifiable Robust Multi-modal Training☆21Jun 21, 2025Updated last year
- Rich Visual Knowledge-based AugmentationNetwork for Visual Question Answering☆10Dec 6, 2019Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official implementation of URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding (AAAI 2026…☆42Feb 4, 2026Updated 6 months ago
- ☆58Jun 8, 2026Updated 2 months ago
- The code of MGCC: Text-based Occluded Person Re-identification via Multi-Granularity Contrastive Consistency Learning☆20Feb 26, 2025Updated last year
- Re-implementation for 'R-VQA: Learning Visual Relation Facts with Semantic Attention for Visual Question Answering'.☆12Mar 13, 2026Updated 5 months ago
- [AAAI24] Official implement of <Beyond Prototypes: Semantic Anchor Regularization for Better Representation Learning>☆23Jan 31, 2024Updated 2 years ago
- FaceShield: Explainable Face Anti-Spoofing with Multimodal Large Language Models☆14Jul 14, 2026Updated last month
- EMNLP 2024 | Style-Specific Neurons for Steering LLMs in Text Style Transfer☆14Mar 23, 2025Updated last year
- The code used to train and run inference with the ColQwen3 model. Welcome to follow and star! ⭐️⭐️⭐️ https://huggingface.co/goodman2001/…☆15Aug 16, 2026Updated last week
- implement n2nmn with pytorch☆19Apr 10, 2019Updated 7 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [CVPR2025] VDocRAG: Retirval-Augmented Generation over Visually-Rich Documents☆67May 26, 2025Updated last year
- Cross-Modal Retrieval with Partially Mismatched Pairs (IEEE TPAMI 2023, PyTorch Code)☆23Sep 17, 2023Updated 2 years ago
- implementation of the paper Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes☆30Mar 24, 2026Updated 5 months ago
- ☆82Jul 31, 2025Updated last year
- [ICCV23] EDAPS: Enhanced Domain-Adaptive Panoptic Segmentation☆32Apr 12, 2024Updated 2 years ago
- [CVPR 2025] Docopilot: Improving Multimodal Models for Document-Level Understanding☆37Jul 22, 2025Updated last year
- Code for [Pattern Recognition] Prompt Learning based Source-free Domain Adaptation for Medical Image Segmentation.☆30Apr 22, 2025Updated last year