A Docker-powered service for PDF document layout analysis. This service provides a powerful and flexible PDF analysis service. The service allows for the segmentation and classification of different parts of PDF pages, identifying the elements such as texts, titles, pictures, tables and so on.
☆1,357Sep 18, 2026Updated last week
Alternatives and similar repositories for pdf-document-layout-analysis
Users that are interested in pdf-document-layout-analysis are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This project aims to extract Table of Contents (TOC) information from PDF files using the outputs generated by the pdf-document-layout-an…☆21Feb 3, 2025Updated last year
- DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception☆2,279Apr 14, 2025Updated last year
- Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical …☆735Updated this week
- python package to parse pdfs with different parsers☆270Sep 12, 2025Updated last year
- ☆1,451Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis☆462Feb 1, 2023Updated 3 years ago
- Toolkit for linearizing PDFs for LLM datasets/training☆19,656Mar 25, 2026Updated 6 months ago
- A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team…☆1,835Mar 17, 2026Updated 6 months ago
- Yet Another Document Translator☆9,604Aug 5, 2026Updated last month
- An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)☆2,086Mar 17, 2026Updated 6 months ago
- Convert PDF to markdown + JSON quickly with high accuracy☆39,974Sep 13, 2026Updated last week
- The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.☆9,051Mar 25, 2026Updated 6 months ago
- A high-quality PDF to Markdown tool based on large language model visual recognition. 一款基于大模型视觉识别的高质量PDF转Markdown工具☆2,273Updated this week
- A Comprehensive Toolkit for High-Quality PDF Content Extraction☆10,030Jan 3, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- OCR, layout analysis, reading order, table recognition in 90+ languages☆21,416Sep 11, 2026Updated 2 weeks ago
- PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.☆6,328Updated this week
- A Faster LayoutReader Model based on LayoutLMv3, Sort OCR bboxes to reading order.☆326Aug 15, 2025Updated last year
- OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex lay…☆2,532Apr 14, 2026Updated 5 months ago
- Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model☆8,221Feb 10, 2025Updated last year
- A lightweight LMM-based Document Parsing Model☆6,652Jul 20, 2026Updated 2 months ago
- ☆16Apr 26, 2024Updated 2 years ago
- Multilingual Document Layout Parsing in a Single Vision-Language Model☆9,152Mar 24, 2026Updated 6 months ago
- Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.☆80,643Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [CVPR 2025] A Comprehensive Benchmark for Document Parsing and Evaluation☆2,065Sep 11, 2026Updated 2 weeks ago
- OCR model that handles complex tables, forms, handwriting with full layout.☆12,328Jun 26, 2026Updated 3 months ago
- Get your documents ready for gen AI☆67,946Updated this week
- ContextGem: Effortless LLM extraction from documents☆2,005Aug 13, 2026Updated last month
- OCR & Document Extraction using vision models☆12,263May 20, 2025Updated last year
- https://no-ocr.com/about☆183Jun 30, 2025Updated last year
- A Repo For Document AI☆3,259Sep 14, 2026Updated last week
- 阅读顺序、Layoutreader☆18May 8, 2025Updated last year
- NeurIPS'24 Workshop - UniTable: Towards a Unified Table Foundation Model☆534Apr 21, 2026Updated 5 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A Unified Toolkit for Deep Learning Based Document Image Analysis☆5,779Aug 15, 2024Updated 2 years ago
- PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.☆10,773Updated this week
- PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.☆29,371Updated this week
- Vision infrastructure to turn complex documents into RAG/LLM-ready data☆4,148Sep 9, 2026Updated 2 weeks ago
- Document Layout Analysis resources repos for development with PdfPig.☆637Oct 1, 2023Updated 2 years ago
- Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.☆1,780Dec 21, 2024Updated last year
- Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/…☆90,207Sep 16, 2026Updated last week