OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex layout handling, complicated table parsing and cross-page content merging.
☆2,532Apr 14, 2026Updated 5 months ago
Alternatives and similar repositories for OCRFlux
Users that are interested in OCRFlux are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A lightweight LMM-based Document Parsing Model☆6,652Jul 20, 2026Updated 2 months ago
- Multilingual Document Layout Parsing in a Single Vision-Language Model☆9,149Mar 24, 2026Updated 6 months ago
- An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)☆2,086Mar 17, 2026Updated 6 months ago
- Toolkit for linearizing PDFs for LLM datasets/training☆19,656Mar 25, 2026Updated 5 months ago
- Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.☆80,592Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.☆9,050Mar 25, 2026Updated 5 months ago
- Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model☆8,219Feb 10, 2025Updated last year
- AI Podcast Generator for bilingual episodes, Multi Languages, Alternative to NotebookLLM;真人对话AI播客生成器,多语言,多音色☆1,291Jul 1, 2025Updated last year
- [CVPR 2025] A Comprehensive Benchmark for Document Parsing and Evaluation☆2,064Sep 11, 2026Updated last week
- ☆1,448Updated this week
- LiYing is an automated photo processing program designed for automating the post-processing workflow of ID photos in general photo studio…☆3,646Aug 9, 2026Updated last month
- OCR, layout analysis, reading order, table recognition in 90+ languages☆21,415Sep 11, 2026Updated last week
- A high-quality PDF to Markdown tool based on large language model visual recognition. 一款基于大模型视觉识别的高质量PDF转Markdown工具☆2,273Updated this week
- 📄 Awesome OCR multiple programing languages toolkits based on ONNX Runtime, OpenVINO, MNN, PaddlePaddle, TensorRT and PyTorch.☆7,938Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/…☆90,158Sep 16, 2026Updated last week
- Convert PDF to markdown + JSON quickly with high accuracy☆39,916Sep 13, 2026Updated last week
- A Comprehensive Toolkit for High-Quality PDF Content Extraction☆10,028Jan 3, 2025Updated last year
- OCR & Document Extraction using vision models☆12,263May 20, 2025Updated last year
- DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception☆2,277Apr 14, 2025Updated last year
- HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better☆2,016Updated this week
- PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.☆6,325Updated this week
- chat log tool, easily use your own chat data. 聊天记录工具,轻松使用自己的聊天数据☆9,187Oct 20, 2025Updated 11 months ago
- [EMNLP 2025 Demo] PDF scientific paper translation with preserved formats - 基于 AI 完整保留排版的 PDF 文档全文双语翻译,支持 Google/DeepL/Ollama/OpenAI 等服务,…☆37,182Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Using GPT to parse PDF☆3,560Apr 17, 2025Updated last year
- A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive vi…☆38,677Updated this week
- Yet Another Document Translator☆9,604Aug 5, 2026Updated last month
- RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to creat…☆91,271Updated this week
- An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System☆24,169Aug 18, 2026Updated last month
- The first open-source agent skills builder. Define skills by vibe workflow, run on Claude Code, Cursor, Codex & more. Build Clawdbot 🦞· …☆7,530Jul 29, 2026Updated last month
- AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs☆52,131Updated this week
- ☆197Dec 7, 2025Updated 9 months ago
- Get your documents ready for gen AI☆67,841Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- OCR software, free and offline. 开源、免费的离线OCR软件。支持截屏/批量导入图片,PDF文档识别,排除水印/页眉页脚,扫描/生成二维码。内置多国语言库。☆47,485Nov 20, 2025Updated 10 months ago
- Offline translation model server with low resource consumption, fast speed, and private deployment capability. 低资源占用速度快可私有部署的离线翻译模型服务器☆4,709Mar 8, 2026Updated 6 months ago
- The official CLI for Agentic Document Extraction (ADE) by LandingAI — parse documents and extract schema-shaped data from your terminal☆2,417Aug 19, 2026Updated last month
- ContextGem: Effortless LLM extraction from documents☆2,004Aug 13, 2026Updated last month
- Tongyi Deep Research, the Leading Open-source Deep Research Agent☆19,988Feb 27, 2026Updated 6 months ago
- Formerly KrillinAI. Open-source AI workspace for creators, powered by Codex. Create videos, images, voice, avatars, video translation, an…☆12,309Updated this week
- An AI agent development platform with all-in-one visual tools, simplifying agent creation, debugging, and deployment like never before. C…☆21,648Jul 29, 2026Updated last month