Simple package to extract text with coordinates from programmatic PDFs
☆333Sep 2, 2026Updated this week
Alternatives and similar repositories for docling-parse
Users that are interested in docling-parse are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆210Updated this week
- Docling core data types and transformations☆281Updated this week
- Create fast graph language models from converted PDF documents for knowledge extraction and Q&A.☆65Updated this week
- Evaluation framework for document processing models and services.☆77Jul 16, 2026Updated last month
- Running Docling as an API service☆1,783Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Making docling agentic through MCP☆732Updated this week
- ☆35Updated this week
- ☆17Apr 8, 2026Updated 4 months ago
- Build document-native LLM applications☆57Sep 11, 2024Updated last year
- Interact with the Deep Search platform for new knowledge explorations and discoveries☆227Jan 24, 2025Updated last year
- Docling LangChain integration☆76Aug 14, 2026Updated 3 weeks ago
- Docling workshops☆43Jul 31, 2026Updated last month
- Get your documents ready for gen AI☆65,896Updated this week
- ☆22Feb 1, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Agent that read, write and edit documents.☆160Aug 19, 2026Updated 2 weeks ago
- Examples using the Deep Search functionalities☆90Jan 29, 2025Updated last year
- Python bindings to PDFium, reasonably cross-platform.☆819Updated this week
- Extract structured text from pdfs quickly☆719Jul 8, 2026Updated last month
- Parallel and LAzY Analyzer for PDFs 🏖️☆47Apr 28, 2026Updated 4 months ago
- This repository provides the code for applying Contrastive Learning Penalty Loss (CLPL) and Mixture of Experts (MoE) to the BGE-M3 text e…☆11Dec 27, 2024Updated last year
- 📚 Process PDFs, Word documents and more with spaCy☆910Mar 27, 2026Updated 5 months ago
- A High-efficiency Open-source Toolkit for Table-to-Latex Task☆276Dec 6, 2025Updated 8 months ago
- Docling Haystack integration☆29Apr 9, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- OnnxTR a docTR (Document Text Recognition) library Onnx pipeline wrapper - for seamless, high-performing & accessible OCR☆195Updated this week
- Open source project for data preparation for GenAI applications☆957Updated this week
- Per-collection OCR leaderboards using VLM-as-judge☆69Aug 24, 2026Updated last week
- LoRA supervised fine-tuning, RLHF (PPO) and RAG with llama-3-8B on the TLDR summarization dataset☆14Feb 2, 2025Updated last year
- Small python package to measure OCR quality and other related metrics.☆26Feb 19, 2024Updated 2 years ago
- Rails application that allows humans to play poker matches managed by the Annual Computer Poker Competition's Dealer program in a web GUI…☆11Apr 25, 2015Updated 11 years ago
- A Docker-powered service for PDF document layout analysis. This service provides a powerful and flexible PDF analysis service. The servic…☆1,349Jul 13, 2026Updated last month
- A Blueprint-style visual node editor for creating FastMCP servers. Build MCP tools, resources, and prompts by connecting nodes - no codin…☆26Dec 8, 2025Updated 8 months ago
- Bajo los adoquines, la PLAYA 🏖️☆17Jul 3, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Convert PDF to markdown + JSON quickly with high accuracy☆39,501Updated this week
- [CVPR 2025] A Comprehensive Benchmark for Document Parsing and Evaluation☆2,015Jul 27, 2026Updated last month
- This repository serves as a collection of scrapers procuring and structuring various legal datasets☆19Jun 16, 2023Updated 3 years ago
- Dataset of PNG images from synthetically generated table layouts with annotations in JSONL files☆154Sep 17, 2025Updated 11 months ago
- ☆17Mar 22, 2024Updated 2 years ago
- A Faster LayoutReader Model based on LayoutLMv3, Sort OCR bboxes to reading order.☆324Aug 15, 2025Updated last year
- PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.☆10,633Updated this week