Multimodal OCR: Parse Anything from Documents
☆302Mar 20, 2026Updated 4 months ago
Alternatives and similar repositories for dots.mocr
Users that are interested in dots.mocr are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Multilingual Document Layout Parsing in a Single Vision-Language Model☆9,016Mar 24, 2026Updated 3 months ago
- (ICCV 2023) ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in Transformer☆78Apr 9, 2024Updated 2 years ago
- Official implementation of URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding (AAAI 2026…☆43Feb 4, 2026Updated 5 months ago
- [Innovation 2026] Oracle bone script decipherment via human-workflow-inspired deep learning☆31Jun 22, 2026Updated 3 weeks ago
- [ICLR 2026] OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning☆76May 26, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Read Ten Lines at One Glance: Line-Aware Semi-Autoregressive Transformer for Multi-Line Handwritten Mathematical Expression Recognition☆28Aug 29, 2023Updated 2 years ago
- Official implementation of ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining (AAAI 20…☆66Jul 4, 2024Updated 2 years ago
- [MM'2024] Official release of RFUND introduced in the MM'2024 paper "PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking f…☆21Dec 4, 2024Updated last year
- This repository is the implementation of "QT-TextSR: Enhancing scene text image super-resolution via efficient interaction with text reco…☆20Jul 9, 2025Updated last year
- ☆1,861Updated this week
- [CVPR 2025] DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding☆30Dec 18, 2025Updated 7 months ago
- [ICCV2025] Training-Free Diffusion Models for Geometric Image Editing☆37Jan 13, 2026Updated 6 months ago
- [CVPR 2025] A Comprehensive Benchmark for Document Parsing and Evaluation☆1,898Jun 26, 2026Updated 3 weeks ago
- ☆289Mar 4, 2026Updated 4 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Official implementation of UPOCR: Towards unified pixel-level OCR interface (ICML 2024)☆69Jun 6, 2024Updated 2 years ago
- VisuRiddles: Fine-grained Perception is a important thing for Multimodal Large Models in Riddles Solving☆20Jun 9, 2026Updated last month
- [AAAI 2025] DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming☆36Jun 1, 2025Updated last year
- DraftClaw is a pre-review tool for academic papers and research documents. Before you submit your draft to reviewers, advisors, or collab…☆19Apr 16, 2026Updated 3 months ago
- DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception☆2,232Apr 14, 2025Updated last year
- Qianfan-VL: Domain-Enhanced Universal Vision-Language Models☆418Mar 18, 2026Updated 4 months ago
- ☆27Jul 5, 2026Updated 2 weeks ago
- MathNet: A Data-Centric Approach, Dataset and Benchmark Model to Advance Mathematical Expression Recognition☆10Mar 19, 2025Updated last year
- [MM'2024] PEneo, an effective algorithm for key-value pair extraction from form-like documents, designed for real-world applications.☆41Apr 7, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆22May 30, 2023Updated 3 years ago
- Dreambooth (LoRA) with well-organized code structure. Naive adaptation from 🤗Diffusers.☆18May 18, 2023Updated 3 years ago
- Official PyTorch implementation of `[ACMMM 2023]Relational Contrastive Learning for Scene Text Recognition`☆17Sep 22, 2023Updated 2 years ago
- [CVPR2026] TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering☆53Updated this week
- [ACL '26] Source code and datasets for our paper "UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Vi…☆17Apr 28, 2026Updated 2 months ago
- The official implement of CTRNet++.☆15Dec 30, 2024Updated last year
- GLM-OCR: Accurate × Fast × Comprehensive☆7,177Apr 21, 2026Updated 2 months ago
- Visual Causal Flow☆3,152Feb 3, 2026Updated 5 months ago
- [AAAI2025 Oral] Predicting the Original Appearance of Damaged Historical Documents☆111Jun 28, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Evaluation of the Optical Character Recognition (OCR) capabilities of GPT-4V(ision)☆128Nov 13, 2023Updated 2 years ago
- OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commer…☆1,415May 20, 2026Updated 2 months ago
- ☆15Nov 26, 2023Updated 2 years ago
- [AAAI 2026 Oral] The official GitHub page of "PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Bas…☆73Jun 22, 2026Updated 3 weeks ago
- RoDLA: Benchmarking the Robustness of Document Layout Analysis Models☆39Mar 26, 2025Updated last year
- [TAI 2025] Official implementation of TAI-accepted paper: ShadowMaskFormer: Mask Augmented Patch Embedding for Shadow Removal☆15May 8, 2025Updated last year
- Turning a CLIP Model into a Scene Text Detector (CVPR2023) | Turning a CLIP Model into a Scene Text Spotter (TPAMI)☆202Jun 17, 2024Updated 2 years ago