python package to parse pdfs with different parsers
☆269Sep 12, 2025Updated 8 months ago
Alternatives and similar repositories for ParseStudio
Users that are interested in ParseStudio are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)☆2,020Mar 17, 2026Updated 2 months ago
- Unlimited text-to-speech in the Browser using Kokoro-JS, 100% local, 100% open source☆345Jun 12, 2025Updated 11 months ago
- POINTS-Reader train☆20Sep 20, 2025Updated 8 months ago
- Fast, zero-copy HTML Parser written in Rust☆28Dec 6, 2025Updated 5 months ago
- A Docker-powered service for PDF document layout analysis. This service provides a powerful and flexible PDF analysis service. The servic…☆1,149May 6, 2026Updated 3 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The minimal, ad-hoc way of plug and play NebulaGraph with pip install, even inside Colab Notebook!☆21May 24, 2024Updated 2 years ago
- A tutorial on DSPy and whether automated prompt engineering lives up to the hype☆26May 3, 2024Updated 2 years ago
- The official repository of NodeRAG☆414Mar 19, 2025Updated last year
- ☆31May 9, 2025Updated last year
- ☆14Mar 4, 2026Updated 2 months ago
- I have explained how to create superior RAG pipeline for complex pdfs using LlamaParse. We can extract text and tables from pdf and QA on…☆48Feb 27, 2024Updated 2 years ago
- A Deep Research agent from scratch☆221May 18, 2025Updated last year
- The Level-Navi Agent, a framework that requires no training and utilizes large language models for deep query understanding and precise s…☆82Dec 27, 2024Updated last year
- Fetch an entire site and use it as an MCP Server☆759Apr 12, 2026Updated last month
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical …☆701May 4, 2026Updated 3 weeks ago
- Beets.io plugin that expose SubSonic API endpoints, allowing you to stream your music everywhere.☆49May 7, 2024Updated 2 years ago
- OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex lay…☆2,511Apr 14, 2026Updated last month
- A lightweight LMM-based Document Parsing Model☆6,593May 8, 2026Updated 3 weeks ago
- (WIP) various language support for libpglite native☆23Aug 5, 2025Updated 9 months ago
- A Comprehensive Toolkit for High-Quality PDF Content Extraction☆9,682Jan 3, 2025Updated last year
- A simple ReAct agent that has access to LlamaIndex docs and to the internet to provide you with insights on LlamaIndex itself.☆11Feb 23, 2025Updated last year
- 影视 分镜大师☆46Nov 27, 2025Updated 6 months ago
- A simple web-based Docker container management interface with a modern design. This application provides a fast and intuitive way to star…☆113Mar 17, 2026Updated 2 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.☆5,666Apr 30, 2026Updated last month
- Customize your arXiv recommendation every day.☆151Sep 24, 2025Updated 8 months ago
- Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pret…☆724Mar 6, 2026Updated 2 months ago
- Autocomplete for contenteditable tags☆12Mar 5, 2021Updated 5 years ago
- Create fast graph language models from converted PDF documents for knowledge extraction and Q&A.☆59Jan 27, 2025Updated last year
- An agentic company research tool powered by LangGraph and Tavily that conducts deep diligence on companies using a multi-agent framework.…☆1,892May 19, 2026Updated last week
- Nebula docker image for development☆16Apr 1, 2026Updated last month
- A high-quality PDF to Markdown tool based on large language model visual recognition. 一款基于大模型视觉识别的高质量PDF转Markdown工具☆1,746Jan 25, 2026Updated 4 months ago
- ☆85Mar 6, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- FlexRAG: A RAG Framework for Information Retrieval and Generation.☆235Apr 28, 2026Updated last month
- A userspace filesystem backing by Apache OpenDAL.☆36Jan 8, 2026Updated 4 months ago
- Fetch arxiv data to LLM-friendly text☆132Feb 18, 2026Updated 3 months ago
- ☆18Apr 7, 2026Updated last month
- Test Environment Booking tool☆14Nov 16, 2020Updated 5 years ago
- [EMNLP 2025] ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents☆660Jan 11, 2026Updated 4 months ago
- Model Context Protocol Server for Apache OpenDAL™☆34Apr 10, 2025Updated last year