OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
☆30Feb 4, 2026Updated 5 months ago
Alternatives and similar repositories for OCRVerse
Users that are interested in OCRVerse are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [CVPR 2026] Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR☆17Mar 23, 2026Updated 4 months ago
- Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner☆24Aug 7, 2025Updated 11 months ago
- ☆21Jan 22, 2026Updated 6 months ago
- ☆67Sep 6, 2025Updated 10 months ago
- ☆42Jan 9, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [EMNLP 2025 Findings] MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation☆15Aug 22, 2025Updated 11 months ago
- Code and Dataset for our paper: Layout-Aware Single-Image Document Flattening☆24Dec 16, 2024Updated last year
- Repository for contributions for Data Generation for Post-OCR correction of Cyrillic handwriting paper☆23Nov 27, 2023Updated 2 years ago
- Learning to Skip the Middle Layers of Transformers☆17Aug 7, 2025Updated 11 months ago
- Squeeze3D: Your 3D Generation Model is Secretly an Extreme Neural Compressor☆23Jun 12, 2025Updated last year
- You found a secret! lzmisscc/lzmisscc is a ✨special ✨ repository that you can use to add a README.md to your GitHub profile. Make sure it…☆13Apr 4, 2026Updated 3 months ago
- [MIR] Pytorch Implementation for FM2S, a denoising algorithm for fluorescence microscopy.☆15Mar 13, 2026Updated 4 months ago
- [ACL 2025] Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQL☆16Oct 9, 2025Updated 9 months ago
- Code for 🌍 UI-Simulator: LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training☆21Oct 17, 2025Updated 9 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- The repository of the ACCV 2024 paper "FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Ge…☆12Jul 28, 2025Updated 11 months ago
- ☆16May 28, 2026Updated last month
- [ICLR 2026] Adaptive Social Learning via Mode Policy Optimization for Language Agents☆51Feb 2, 2026Updated 5 months ago
- The official implementation of BFSM.☆18Sep 30, 2025Updated 9 months ago
- The public reproducible analysis code used for the gaze project☆11May 16, 2026Updated 2 months ago
- [ECCV 2024] Official PyTorch implementation of LUT "Learning with Unmasked Tokens Drives Stronger Vision Learners"☆14Dec 1, 2024Updated last year
- ☆15Apr 8, 2026Updated 3 months ago
- a model zoo☆11Jul 19, 2017Updated 9 years ago
- Code for paper: Unified Text-to-Image Generation and Retrieval☆16Jul 19, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ACL 2025 Oral] The official repository of our paper: CADReview: Automatically Reviewing CAD Programs with Error Detection and Correction☆23Aug 8, 2025Updated 11 months ago
- Official Implementation for the paper "VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models"☆23Aug 14, 2025Updated 11 months ago
- [ICCV 2023] Efficient Video Action Detection with Token Dropout and Context Refinement☆39Sep 27, 2023Updated 2 years ago
- Code and benchmark for the paper: "A Practitioner's Guide to Continual Multimodal Pretraining" [NeurIPS'24]☆62Dec 10, 2024Updated last year
- [ICLR2025 Oral] ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding☆101Apr 1, 2025Updated last year
- ☆27Feb 27, 2026Updated 5 months ago
- The code for paper: Decoupled Planning and Execution: A Hierarchical Reasoning Framework for Deep Search [SIGIR 2026]☆65Jul 4, 2025Updated last year
- ☆16Apr 10, 2025Updated last year
- Youtu-Parsing: Perception, Structuring and Recognition via High-Parallelism Decoding☆69Jun 15, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- 中文论文、证券类、财报类PDF数据☆41Jun 13, 2024Updated 2 years ago
- Official github repo for "Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute"☆17Jun 30, 2025Updated last year
- 🔍📃 LLM-powered PDF Table Extractor☆19Jun 26, 2025Updated last year
- CAD - Memory Efficient Convolutional Adapter for Segment Anything☆12Oct 4, 2024Updated last year
- The implementation of our NeurIPS 2024 paper "DarkSAM: Fooling Segment Anything Model to Segment Nothing".☆14Nov 4, 2024Updated last year
- LLaVA combines with Magvit Image tokenizer, training MLLM without an Vision Encoder. Unifying image understanding and generation.☆38Jun 20, 2024Updated 2 years ago
- Code for "Saliency Prediction of Sports Videos: A Large-Scale Database and a Self-Adaptive Approach", ICASSP 2024☆14May 28, 2024Updated 2 years ago