Open-source toolkit for reliable RAG pipelines: convert PDFs to Markdown, clean documents, inspect chunks, compare chunking strategies, and enrich metadata for LLM applications.
☆150Jul 7, 2026Updated 2 weeks ago
Alternatives and similar repositories for chunky
Users that are interested in chunky are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Evaluate and compare chunking strategies for RAG pipelines☆42May 30, 2026Updated last month
- Dasein managed vector index service - Python SDK☆20May 11, 2026Updated 2 months ago
- A premium RAG-based AI Assistant built with React and FastAPI. Features efficient document indexing and high-accuracy retrieval-augmented…☆18Jun 3, 2026Updated last month
- Declarative Document Indexing (DDI) framework for Python. Define schemas, extract structured indices, search smarter.☆46Jun 22, 2026Updated 3 weeks ago
- A local-first, self-organizing AI RAG GRAPH knowledge system that reads, links and reasons over your documents — offline by default, clou…☆26Apr 22, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Type-safe AI agents for Go. Suricata combines LLM intelligence with Go’s strong typing, declarative YAML specs, and code generation to bu…☆21Aug 19, 2025Updated 11 months ago
- Official Detectron2 implementation of STMDA-RetinaNet, A Multi Camera Unsupervised Domain Adaptation Pipeline for Object Detection in Cul…☆29Dec 3, 2023Updated 2 years ago
- Documentation☆236Jun 27, 2026Updated 3 weeks ago
- RAG for financial Analysis using SQL and calculator tool.☆16Apr 29, 2026Updated 2 months ago
- A fast and strongly encrypted key-value store written in pure Golang.☆14Jan 30, 2022Updated 4 years ago
- A command line tool for searching and downloading files from the IRC network.☆42Jan 4, 2024Updated 2 years ago
- CDRAG is a new retrieval framework that uses hierarchical document clustering, and LLM-guided document selection from those clusters to c…☆37Apr 13, 2026Updated 3 months ago
- File indexer with semantic search, hybrid retrieval, and multi-step reasoning agents☆21Jan 17, 2026Updated 6 months ago
- Tree-based, vectorless document RAG framework. Connect any LLM via URL/API key.☆40Apr 7, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆13Mar 28, 2024Updated 2 years ago
- Ask questions and get answers from earnings calls, SEC filings and news☆137Jun 19, 2026Updated last month
- Turn scattered knowledge, operational data, and history into source-linked context that your agents can inspect, explain, and reuse.☆42Jun 17, 2026Updated last month
- Roaring Bitmap positional phrase matching for low-latency LLM context retrieval.☆29May 4, 2026Updated 2 months ago
- A Docker-powered RAG system that understands the difference between code and prose. Ingest your codebase and documentation, then query th…☆351Mar 14, 2026Updated 4 months ago
- Find your files with natural language and ask questions.☆63Jul 10, 2026Updated last week
- PDFStract - Extract, Chunking and Embedding Layer in Your RAG Pipeline - Available as CLI - WEBUI - API☆153Mar 18, 2026Updated 4 months ago
- Privacy-first document intelligence engine — parse PDFs, DOCX, PPTX, XLSX & CSV into AI-ready chunks for RAG pipelines. Includes HITL rev…☆30May 5, 2026Updated 2 months ago
- ☆16Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A structured enforcement system for AI-generated code. Turns predictable AI failure modes into architectural gates, static analysis halt…☆24Mar 20, 2026Updated 4 months ago
- ☆10Jul 20, 2023Updated 3 years ago
- Detectron2 implementation of DA-Faster R-CNN, Domain Adaptive Faster R-CNN for Object Detection in the Wild, Computer Vision and Pattern …☆64Jan 21, 2026Updated 6 months ago
- A docs-first LangGraph cookbook: routing, RAG gating, tool-first overrides, and memory.☆20Mar 5, 2026Updated 4 months ago
- FastMemory is a topological representation of text data using concepts as the primary input. It helps in improving the RAG(by replacing e…☆53Jun 8, 2026Updated last month
- Multi-LLM agent framework with Claude Code-like tools. Use DeepSeek, Claude, GPT, Llama, or any model — same tools, same skills, swap fre…☆18Apr 1, 2026Updated 3 months ago
- 🦛 CHONK docs with Chonkie ✨ — The lightweight ingestion library for fast, efficient and robust RAG pipelines☆4,544Updated this week
- MODE (Mixture of Document Experts) is an advanced RAG framework that enhances query response quality by combining hierarchical document c…☆16Sep 3, 2025Updated 10 months ago
- weighted category-balanced dataset builder for LLM fine-tuning☆16Feb 21, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- An opinionated list of practical tools for Conceptual Modeling and Linked Data☆40Mar 24, 2026Updated 3 months ago
- Datacore/React AI Prompt Library for Obsidian, easily plug and play within your vault.☆15May 26, 2025Updated last year
- pdfLLM is a completely open source, proof of concept RAG app.☆186Sep 1, 2025Updated 10 months ago
- Retrieval-augmented generation (RAG) for remote & local LLM use☆45May 24, 2025Updated last year
- [H] HyperspaceDB is a high-performance, vector database. It features 1-bit quantization, async replication, and native support for hierar…☆144Jul 14, 2026Updated last week
- A lightweight cron job scheduler with webhook notifications, designed for easy task automation and flexible scheduling☆58Jan 11, 2025Updated last year
- I developed a fine-tuned retrieval head for RAG that learns to more reliably retrieve relevant passages by transforming the query embeddi…☆17May 21, 2026Updated 2 months ago