python package to parse pdfs with different parsers
☆269Sep 12, 2025Updated 9 months ago
Alternatives and similar repositories for ParseStudio
Users that are interested in ParseStudio are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)☆2,025Mar 17, 2026Updated 3 months ago
- Unlimited text-to-speech in the Browser using Kokoro-JS, 100% local, 100% open source☆344Jun 12, 2025Updated last year
- A Docker-powered service for PDF document layout analysis. This service provides a powerful and flexible PDF analysis service. The servic…☆1,152May 6, 2026Updated last month
- Parallel and LAzY Analyzer for PDFs 🏖️☆45Apr 28, 2026Updated last month
- The official repository of NodeRAG☆416Mar 19, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- ☆31May 9, 2025Updated last year
- ☆14Mar 4, 2026Updated 3 months ago
- I have explained how to create superior RAG pipeline for complex pdfs using LlamaParse. We can extract text and tables from pdf and QA on…☆48Feb 27, 2024Updated 2 years ago
- The Level-Navi Agent, a framework that requires no training and utilizes large language models for deep query understanding and precise s…☆82Dec 27, 2024Updated last year
- Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical …☆712May 4, 2026Updated last month
- OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex lay…☆2,512Apr 14, 2026Updated 2 months ago
- A Comprehensive Toolkit for High-Quality PDF Content Extraction☆9,728Jan 3, 2025Updated last year
- A simple ReAct agent that has access to LlamaIndex docs and to the internet to provide you with insights on LlamaIndex itself.☆11Feb 23, 2025Updated last year
- A simple web-based Docker container management interface with a modern design. This application provides a fast and intuitive way to star…☆114Mar 17, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.☆5,759Jun 6, 2026Updated last week
- ☆22Feb 1, 2025Updated last year
- Customize your arXiv recommendation every day.☆152Sep 24, 2025Updated 8 months ago
- Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pret…☆724Mar 6, 2026Updated 3 months ago
- Create fast graph language models from converted PDF documents for knowledge extraction and Q&A.☆59Jan 27, 2025Updated last year
- An agentic company research tool powered by LangGraph and Tavily that conducts deep diligence on companies using a multi-agent framework.…☆1,967May 19, 2026Updated last month
- Nebula docker image for development☆16Apr 1, 2026Updated 2 months ago
- A high-quality PDF to Markdown tool based on large language model visual recognition. 一款基于大模型视觉识别的高质量PDF转Markdown工具☆1,758Jan 25, 2026Updated 4 months ago
- ☆88Mar 6, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The official repository of "Document Image Machine Translation with Dynamic Multi-pre-trained Models Assembling"☆14Nov 26, 2025Updated 6 months ago
- FlexRAG: A RAG Framework for Information Retrieval and Generation.☆237Jun 3, 2026Updated 2 weeks ago
- A userspace filesystem backing by Apache OpenDAL.☆38Jun 2, 2026Updated 2 weeks ago
- Fetch arxiv data to LLM-friendly text☆132Feb 18, 2026Updated 4 months ago
- Minimalist agent framework for AI engineers☆21Jan 22, 2026Updated 4 months ago
- Simple package to extract text with coordinates from programmatic PDFs☆305Jun 1, 2026Updated 2 weeks ago
- Test Environment Booking tool☆14Nov 16, 2020Updated 5 years ago
- [EMNLP 2025] ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents☆664Jan 11, 2026Updated 5 months ago
- Model Context Protocol Server for Apache OpenDAL™☆34Apr 10, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- coze api to openai☆15Sep 1, 2024Updated last year
- ☆48Dec 4, 2025Updated 6 months ago
- A minimal Openclaw built using the Opencode SDK☆82Feb 7, 2026Updated 4 months ago
- 小智ai机器人☆10Mar 8, 2025Updated last year
- Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval And Synthesis For SLMs☆59Oct 7, 2025Updated 8 months ago
- ☆273Nov 15, 2024Updated last year
- A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines☆5,583Jun 12, 2026Updated last week