Open-source toolkit for reliable RAG pipelines: convert PDFs to Markdown, clean documents, inspect chunks, compare chunking strategies, and enrich metadata for LLM applications.
☆180Aug 30, 2026Updated 3 weeks ago
Alternatives and similar repositories for chunky
Users that are interested in chunky are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Evaluate and compare chunking strategies for RAG pipelines☆42May 30, 2026Updated 3 months ago
- Dasein managed vector index service - Python SDK☆23May 11, 2026Updated 4 months ago
- A modular Agentic RAG built with LangGraph — learn Retrieval-Augmented Generation Agents in minutes.☆4,194Aug 30, 2026Updated 3 weeks ago
- Declarative Document Indexing (DDI) framework for Python. Define schemas, extract structured indices, search smarter.☆47Jun 22, 2026Updated 3 months ago
- Visual document analysis studio powered by Docling — configure the extraction pipeline, inspect text, tables and bounding boxes in the br…☆261Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- RAG for financial Analysis using SQL and calculator tool.☆15Jul 25, 2026Updated last month
- A Flask and bootsrap based hospital managemnt system to help manage and monitor the basic day today activities of a hospital, complete wi…☆10Jun 29, 2021Updated 5 years ago
- Ask questions and get answers from earnings calls, SEC filings and news☆143Jun 19, 2026Updated 3 months ago
- Experimental RAG playground for exploring retrieval quality, corpus construction, and filter-chain design. Features configurable ranking …☆30Updated this week
- Turn scattered knowledge, operational data, and history into source-linked context that your agents can inspect, explain, and reuse.☆58Jun 17, 2026Updated 3 months ago
- Roaring Bitmap positional phrase matching for low-latency LLM context retrieval.☆29May 4, 2026Updated 4 months ago
- A Docker-powered RAG system that understands the difference between code and prose. Ingest your codebase and documentation, then query th…☆354Mar 14, 2026Updated 6 months ago
- Find your files with natural language and ask questions.☆64Aug 21, 2026Updated last month
- PDFStract - Extract, Chunking and Embedding Layer in Your RAG Pipeline - Available as CLI - WEBUI - API☆153Mar 18, 2026Updated 6 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ✨基于Django的学生管理系统,数据库MySQL或其他,前端Bootstrap(想要✨star!)☆12Apr 13, 2023Updated 3 years ago
- ☆16Aug 10, 2026Updated last month
- ☆10Jul 20, 2023Updated 3 years ago
- A docs-first LangGraph cookbook: routing, RAG gating, tool-first overrides, and memory.☆20Mar 5, 2026Updated 6 months ago
- FastMemory is a topological representation of text data using concepts as the primary input. It helps in improving the RAG(by replacing e…☆53Jun 8, 2026Updated 3 months ago
- Multi-LLM agent framework with Claude Code-like tools. Use DeepSeek, Claude, GPT, Llama, or any model — same tools, same skills, swap fre…☆18Apr 1, 2026Updated 5 months ago
- Plugin for interacting with LLMs in Polars☆30Updated this week
- ☆17Oct 6, 2025Updated 11 months ago
- 🦛 CHONK docs with Chonkie ✨ — The lightweight ingestion library for fast, efficient and robust RAG pipelines☆4,766Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- One library to split them all: Sentence, Code, Docs. Chunk smarter, not harde; built for LLMs, RAG pipelines, and beyond.☆84Updated this week
- MODE (Mixture of Document Experts) is an advanced RAG framework that enhances query response quality by combining hierarchical document c…☆17Sep 3, 2025Updated last year
- weighted category-balanced dataset builder for LLM fine-tuning☆17Feb 21, 2026Updated 7 months ago
- pdfLLM is a completely open source, proof of concept RAG app.☆187Sep 1, 2025Updated last year
- A docs-first guide to LLM system design — hybrid search, embedding pipelines, reranking, and LLM-as-judge patterns.☆45Jun 22, 2026Updated 3 months ago
- Hybrid RAG system combining vector search, knowledge graph (LightRAG), and cross-encoder reranking — with Docling document parsing, visua…☆525Apr 20, 2026Updated 5 months ago
- Retrieval-augmented generation (RAG) for remote & local LLM use☆45May 24, 2025Updated last year
- I developed a fine-tuned retrieval head for RAG that learns to more reliably retrieve relevant passages by transforming the query embeddi…☆17May 21, 2026Updated 4 months ago
- This is more like a QA system☆16Jul 20, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Exploring retrieval systems for language models☆14Aug 7, 2026Updated last month
- Implementation of a fast semantic chunker in C++, installable in python 3.7+ projects.☆22Sep 20, 2025Updated last year
- Geometric AI research: a proven cube-math core, reusable vector-collapse dynamics, and reproducible experiments in embeddings, NLI, gener…☆15Jul 31, 2026Updated last month
- Fully local governed RAG for code, documents, and tables. No LLM API key. No GPU. No infrastructure.☆41Updated this week
- Dataset and benchmark for RAG on company internal documents.☆566Sep 3, 2026Updated 2 weeks ago
- Docling simplifies document processing, parsing diverse formats — including advanced PDF understanding — and providing seamless integrati…☆23Updated this week
- A Claude skill that builds PPTX decks matching your real designer decks: extracted design DNA, measured word budget, render-QA loop.☆42Jul 11, 2026Updated 2 months ago