A Docker-powered service for PDF document layout analysis. This service provides a powerful and flexible PDF analysis service. The service allows for the segmentation and classification of different parts of PDF pages, identifying the elements such as texts, titles, pictures, tables and so on.
☆1,350Jul 13, 2026Updated last month
Alternatives and similar repositories for pdf-document-layout-analysis
Users that are interested in pdf-document-layout-analysis are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This project aims to extract Table of Contents (TOC) information from PDF files using the outputs generated by the pdf-document-layout-an…☆21Feb 3, 2025Updated last year
- DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception☆2,263Apr 14, 2025Updated last year
- Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical …☆733Updated this week
- python package to parse pdfs with different parsers☆269Sep 12, 2025Updated 11 months ago
- ☆1,405May 13, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis☆458Feb 1, 2023Updated 3 years ago
- Toolkit for linearizing PDFs for LLM datasets/training☆19,437Mar 25, 2026Updated 5 months ago
- A collection of original, innovative ideas and algorithms towards Advanced Literate Machinery. This project is maintained by the OCR Team…☆1,835Mar 17, 2026Updated 5 months ago
- Yet Another Document Translator☆9,477Aug 5, 2026Updated last month
- An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)☆2,086Mar 17, 2026Updated 5 months ago
- Convert PDF to markdown + JSON quickly with high accuracy☆39,531Updated this week
- The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.☆9,050Mar 25, 2026Updated 5 months ago
- A high-quality PDF to Markdown tool based on large language model visual recognition. 一款基于大模型视觉识别的高质量PDF转Markdown工具☆2,193Jan 25, 2026Updated 7 months ago
- A Comprehensive Toolkit for High-Quality PDF Content Extraction☆10,008Jan 3, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- OCR, layout analysis, reading order, table recognition in 90+ languages☆21,352Updated this week
- PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.☆6,271Updated this week
- A Faster LayoutReader Model based on LayoutLMv3, Sort OCR bboxes to reading order.☆324Aug 15, 2025Updated last year
- OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex lay…☆2,532Apr 14, 2026Updated 4 months ago
- Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model☆8,221Feb 10, 2025Updated last year
- A lightweight LMM-based Document Parsing Model☆6,636Jul 20, 2026Updated last month
- ☆16Apr 26, 2024Updated 2 years ago
- Multilingual Document Layout Parsing in a Single Vision-Language Model☆9,104Mar 24, 2026Updated 5 months ago
- Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.☆79,250Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [CVPR 2025] A Comprehensive Benchmark for Document Parsing and Evaluation☆2,019Jul 27, 2026Updated last month
- OCR model that handles complex tables, forms, handwriting with full layout.☆12,221Jun 26, 2026Updated 2 months ago
- Get your documents ready for gen AI☆66,035Updated this week
- ContextGem: Effortless LLM extraction from documents☆1,997Aug 13, 2026Updated 3 weeks ago
- OCR & Document Extraction using vision models☆12,266May 20, 2025Updated last year
- https://no-ocr.com/about☆182Jun 30, 2025Updated last year
- A Repo For Document AI☆3,257Updated this week
- 阅读顺序、Layoutreader☆18May 8, 2025Updated last year
- NeurIPS'24 Workshop - UniTable: Towards a Unified Table Foundation Model☆533Apr 21, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A Unified Toolkit for Deep Learning Based Document Image Analysis☆5,778Aug 15, 2024Updated 2 years ago
- PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.☆10,642Updated this week
- PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.☆28,949Updated this week
- Vision infrastructure to turn complex documents into RAG/LLM-ready data☆4,140Apr 9, 2026Updated 4 months ago
- Document Layout Analysis resources repos for development with PdfPig.☆636Oct 1, 2023Updated 2 years ago
- Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.☆1,772Dec 21, 2024Updated last year
- Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/…☆88,927Jul 22, 2026Updated last month