python package to parse pdfs with different parsers
☆250Sep 12, 2025Updated 7 months ago
Alternatives and similar repositories for ParseStudio
Users that are interested in ParseStudio are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)☆1,968Mar 17, 2026Updated last month
- Unlimited text-to-speech in the Browser using Kokoro-JS, 100% local, 100% open source☆336Jun 12, 2025Updated 10 months ago
- Fast, zero-copy HTML Parser written in Rust☆27Dec 6, 2025Updated 5 months ago
- A Docker-powered service for PDF document layout analysis. This service provides a powerful and flexible PDF analysis service. The servic…☆1,113Updated this week
- The minimal, ad-hoc way of plug and play NebulaGraph with pip install, even inside Colab Notebook!☆21May 24, 2024Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Parallel and LAzY Analyzer for PDFs 🏖️☆41Apr 28, 2026Updated last week
- The official repository of NodeRAG☆412Mar 19, 2025Updated last year
- ☆30May 9, 2025Updated last year
- ☆14Mar 4, 2026Updated 2 months ago
- I have explained how to create superior RAG pipeline for complex pdfs using LlamaParse. We can extract text and tables from pdf and QA on…☆48Feb 27, 2024Updated 2 years ago
- A Deep Research agent from scratch☆219May 18, 2025Updated 11 months ago
- The Level-Navi Agent, a framework that requires no training and utilizes large language models for deep query understanding and precise s…☆82Dec 27, 2024Updated last year
- A set of tools to create synthetically-generated data from documents☆45Aug 15, 2025Updated 8 months ago
- Fetch an entire site and use it as an MCP Server☆759Apr 12, 2026Updated 3 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical …☆661Updated this week
- 新 React 中文文档 docker 版本,方便本地部署看文档☆13Apr 23, 2026Updated 2 weeks ago
- OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex lay…☆2,494Apr 14, 2026Updated 3 weeks ago
- A lightweight LMM-based Document Parsing Model☆6,594Updated this week
- (WIP) various language support for libpglite native☆22Aug 5, 2025Updated 9 months ago
- A Comprehensive Toolkit for High-Quality PDF Content Extraction☆9,642Jan 3, 2025Updated last year
- A simple ReAct agent that has access to LlamaIndex docs and to the internet to provide you with insights on LlamaIndex itself.☆11Feb 23, 2025Updated last year
- A simple web-based Docker container management interface with a modern design. This application provides a fast and intuitive way to star…☆114Mar 17, 2026Updated last month
- ☆22Feb 1, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Customize your arXiv recommendation every day.☆150Sep 24, 2025Updated 7 months ago
- Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pret…☆722Mar 6, 2026Updated 2 months ago
- PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.☆5,488Apr 30, 2026Updated last week
- An agentic company research tool powered by LangGraph and Tavily that conducts deep diligence on companies using a multi-agent framework.…☆1,883Updated this week
- Nebula docker image for development☆16Apr 1, 2026Updated last month
- ☆83Mar 6, 2026Updated 2 months ago
- The official repository of "Document Image Machine Translation with Dynamic Multi-pre-trained Models Assembling"☆14Nov 26, 2025Updated 5 months ago
- FlexRAG: A RAG Framework for Information Retrieval and Generation.☆236Apr 28, 2026Updated last week
- A userspace filesystem backing by Apache OpenDAL.☆36Jan 8, 2026Updated 4 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Fetch arxiv data to LLM-friendly text☆132Feb 18, 2026Updated 2 months ago
- Simple package to extract text with coordinates from programmatic PDFs☆272Updated this week
- Test Environment Booking tool☆14Nov 16, 2020Updated 5 years ago
- [EMNLP 2025] ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents☆658Jan 11, 2026Updated 3 months ago
- Model Context Protocol Server for Apache OpenDAL™☆34Apr 10, 2025Updated last year
- A minimal Openclaw built using the Opencode SDK☆81Feb 7, 2026Updated 3 months ago
- 小智ai机器人☆10Mar 8, 2025Updated last year