A Docker-powered service for PDF document layout analysis. This service provides a powerful and flexible PDF analysis service. The service allows for the segmentation and classification of different parts of PDF pages, identifying the elements such as texts, titles, pictures, tables and so on.
☆1,337Jul 13, 2026Updated last month
Alternatives and similar repositories for pdf-document-layout-analysis
Users that are interested in pdf-document-layout-analysis are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This project aims to extract Table of Contents (TOC) information from PDF files using the outputs generated by the pdf-document-layout-an…☆21Feb 3, 2025Updated last year
- DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception☆2,248Apr 14, 2025Updated last year
- Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical …☆719Updated this week
- python package to parse pdfs with different parsers☆269Sep 12, 2025Updated 11 months ago
- ☆1,400May 13, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis☆453Feb 1, 2023Updated 3 years ago
- Toolkit for linearizing PDFs for LLM datasets/training☆19,328Mar 25, 2026Updated 4 months ago
- A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team…☆1,834Mar 17, 2026Updated 5 months ago
- Yet Another Document Translator☆9,349Aug 5, 2026Updated last week
- An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)☆2,038Mar 17, 2026Updated 4 months ago
- Convert PDF to markdown + JSON quickly with high accuracy☆38,781Aug 7, 2026Updated last week
- The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.☆9,047Mar 25, 2026Updated 4 months ago
- A high-quality PDF to Markdown tool based on large language model visual recognition. 一款基于大模型视觉识别的高质量PDF转Markdown工具☆2,171Jan 25, 2026Updated 6 months ago
- A Comprehensive Toolkit for High-Quality PDF Content Extraction☆9,954Jan 3, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- OCR, layout analysis, reading order, table recognition in 90+ languages☆21,286Jul 23, 2026Updated 3 weeks ago
- PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.☆6,130Updated this week
- A Faster LayoutReader Model based on LayoutLMv3, Sort OCR bboxes to reading order.☆324Aug 15, 2025Updated last year
- OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex lay…☆2,533Apr 14, 2026Updated 4 months ago
- Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model☆8,214Feb 10, 2025Updated last year
- A lightweight LMM-based Document Parsing Model☆6,625Jul 20, 2026Updated 3 weeks ago
- ☆16Apr 26, 2024Updated 2 years ago
- Multilingual Document Layout Parsing in a Single Vision-Language Model☆9,071Mar 24, 2026Updated 4 months ago
- Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.☆77,735Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [CVPR 2025] A Comprehensive Benchmark for Document Parsing and Evaluation☆1,975Jul 27, 2026Updated 2 weeks ago
- OCR model that handles complex tables, forms, handwriting with full layout.☆12,054Jun 26, 2026Updated last month
- Get your documents ready for gen AI☆64,836Updated this week
- ContextGem: Effortless LLM extraction from documents☆1,984Updated this week
- OCR & Document Extraction using vision models☆12,266May 20, 2025Updated last year
- https://no-ocr.com/about☆182Jun 30, 2025Updated last year
- A Repo For Document AI☆3,216Updated this week
- UniTable: Towards a Unified Table Foundation Model☆534Apr 21, 2026Updated 3 months ago
- 阅读顺序、Layoutreader☆18May 8, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A Unified Toolkit for Deep Learning Based Document Image Analysis☆5,771Aug 15, 2024Updated 2 years ago
- PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.☆10,485Updated this week
- PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.☆28,431Updated this week
- Vision infrastructure to turn complex documents into RAG/LLM-ready data☆4,133Apr 9, 2026Updated 4 months ago
- Document Layout Analysis resources repos for development with PdfPig.☆637Oct 1, 2023Updated 2 years ago
- Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.☆1,771Dec 21, 2024Updated last year
- Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/…☆87,738Jul 22, 2026Updated 3 weeks ago