python package to parse pdfs with different parsers
☆249Sep 12, 2025Updated 6 months ago
Alternatives and similar repositories for ParseStudio
Users that are interested in ParseStudio are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)☆1,936Mar 17, 2026Updated last week
- Unlimited text-to-speech in the Browser using Kokoro-JS, 100% local, 100% open source☆331Jun 12, 2025Updated 9 months ago
- Github implementation of https://reports.chatclimate.ai/☆23Jun 16, 2025Updated 9 months ago
- POINTS-Reader train☆20Sep 20, 2025Updated 6 months ago
- A Docker-powered service for PDF document layout analysis. This service provides a powerful and flexible PDF analysis service. The servic…☆1,099Mar 2, 2026Updated 3 weeks ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- The minimal, ad-hoc way of plug and play NebulaGraph with pip install, even inside Colab Notebook!☆21May 24, 2024Updated last year
- ☆24Nov 29, 2017Updated 8 years ago
- Parallel and LAzY Analyzer for PDFs 🏖️☆40Mar 9, 2026Updated 2 weeks ago
- A tutorial on DSPy and whether automated prompt engineering lives up to the hype☆26May 3, 2024Updated last year
- ☆14Mar 4, 2026Updated 3 weeks ago
- The official repository of NodeRAG☆410Mar 19, 2025Updated last year
- ☆30May 9, 2025Updated 10 months ago
- I have explained how to create superior RAG pipeline for complex pdfs using LlamaParse. We can extract text and tables from pdf and QA on…☆49Feb 27, 2024Updated 2 years ago
- A Deep Research agent from scratch☆219May 18, 2025Updated 10 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- The Level-Navi Agent, a framework that requires no training and utilizes large language models for deep query understanding and precise s…☆82Dec 27, 2024Updated last year
- Fetch an entire site and use it as an MCP Server☆755Nov 24, 2025Updated 4 months ago
- A set of tools to create synthetically-generated data from documents☆44Aug 15, 2025Updated 7 months ago
- Dedoc is a library (service) for automate documents parsing and bringing to a uniform format. It automatically extracts content, logical …☆653Mar 19, 2026Updated last week
- OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex lay…☆2,482Aug 4, 2025Updated 7 months ago
- 影视分镜大师☆46Nov 27, 2025Updated 4 months ago
- A lightweight LMM-based Document Parsing Model☆6,573Updated this week
- (WIP) various language support for libpglite native☆21Aug 5, 2025Updated 7 months ago
- A Comprehensive Toolkit for High-Quality PDF Content Extraction☆9,503Jan 3, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- A simple ReAct agent that has access to LlamaIndex docs and to the internet to provide you with insights on LlamaIndex itself.☆11Feb 23, 2025Updated last year
- A simple web-based Docker container management interface with a modern design. This application provides a fast and intuitive way to star…☆113Mar 17, 2026Updated last week
- An agentic company research tool powered by LangGraph and Tavily that conducts deep diligence on companies using a multi-agent framework.…☆1,644Updated this week
- ☆22Feb 1, 2025Updated last year
- Customize your arXiv recommendation every day.☆144Sep 24, 2025Updated 6 months ago
- Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pret…☆718Mar 6, 2026Updated 3 weeks ago
- ☆79Mar 6, 2026Updated 3 weeks ago
- PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.☆5,203Updated this week
- Nebula docker image for development☆16Mar 20, 2026Updated last week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- The official repository of "Document Image Machine Translation with Dynamic Multi-pre-trained Models Assembling"☆14Nov 26, 2025Updated 4 months ago
- FlexRAG: A RAG Framework for Information Retrieval and Generation.☆235Updated this week
- Fetch arxiv data to LLM-friendly text☆131Feb 18, 2026Updated last month
- Simple package to extract text with coordinates from programmatic PDFs☆260Updated this week
- A userspace filesystem backing by Apache OpenDAL.☆37Jan 8, 2026Updated 2 months ago
- ☆15Updated this week
- Test Environment Booking tool☆14Nov 16, 2020Updated 5 years ago