Open-source toolkit for reliable RAG pipelines: convert PDFs to Markdown, clean documents, inspect chunks, compare chunking strategies, and enrich metadata for LLM applications.
☆160Jul 25, 2026Updated 2 weeks ago
Alternatives and similar repositories for chunky
Users that are interested in chunky are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Evaluate and compare chunking strategies for RAG pipelines☆42May 30, 2026Updated 2 months ago
- Dasein managed vector index service - Python SDK☆22May 11, 2026Updated 2 months ago
- A modular Agentic RAG built with LangGraph — learn Retrieval-Augmented Generation Agents in minutes.☆3,877Jul 25, 2026Updated 2 weeks ago
- Declarative Document Indexing (DDI) framework for Python. Define schemas, extract structured indices, search smarter.☆47Jun 22, 2026Updated last month
- Type-safe AI agents for Go. Suricata combines LLM intelligence with Go’s strong typing, declarative YAML specs, and code generation to bu…☆21Aug 19, 2025Updated 11 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Documentation☆243Updated this week
- RAG for financial Analysis using SQL and calculator tool.☆15Jul 25, 2026Updated 2 weeks ago
- A fast and strongly encrypted key-value store written in pure Golang.☆14Jan 30, 2022Updated 4 years ago
- CDRAG is a new retrieval framework that uses hierarchical document clustering, and LLM-guided document selection from those clusters to c…☆37Apr 13, 2026Updated 3 months ago
- File indexer with semantic search, hybrid retrieval, and multi-step reasoning agents☆21Jan 17, 2026Updated 6 months ago
- Official Detectron2 implementation of DA-RetinaNet, An unsupervised domain adaptation scheme for single-stage artwork recognition in cult…☆65Dec 3, 2023Updated 2 years ago
- Tree-based, vectorless document RAG framework. Connect any LLM via URL/API key.☆41Apr 7, 2026Updated 4 months ago
- Ask questions and get answers from earnings calls, SEC filings and news☆138Jun 19, 2026Updated last month
- Turn scattered knowledge, operational data, and history into source-linked context that your agents can inspect, explain, and reuse.☆45Jun 17, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Experimental RAG playground for exploring retrieval quality, corpus construction, and filter-chain design. Features configurable ranking …☆27Updated this week
- Roaring Bitmap positional phrase matching for low-latency LLM context retrieval.☆29May 4, 2026Updated 3 months ago
- A Docker-powered RAG system that understands the difference between code and prose. Ingest your codebase and documentation, then query th…☆351Mar 14, 2026Updated 4 months ago
- PDFStract - Extract, Chunking and Embedding Layer in Your RAG Pipeline - Available as CLI - WEBUI - API☆152Mar 18, 2026Updated 4 months ago
- Privacy-first document intelligence engine — parse PDFs, DOCX, PPTX, XLSX & CSV into AI-ready chunks for RAG pipelines. Includes HITL rev…☆30May 5, 2026Updated 3 months ago
- ☆16Jul 19, 2026Updated 3 weeks ago
- ☆10Jul 20, 2023Updated 3 years ago
- FastMemory is a topological representation of text data using concepts as the primary input. It helps in improving the RAG(by replacing e…☆53Jun 8, 2026Updated 2 months ago
- A docs-first LangGraph cookbook: routing, RAG gating, tool-first overrides, and memory.☆20Mar 5, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Multi-LLM agent framework with Claude Code-like tools. Use DeepSeek, Claude, GPT, Llama, or any model — same tools, same skills, swap fre…☆18Apr 1, 2026Updated 4 months ago
- Plugin for interacting with LLMs in Polars☆30Aug 3, 2026Updated last week
- Use any AI model, locally or in the cloud, while tracking on-device token usage and collecting token telemetry data.☆18Jul 31, 2026Updated last week
- ULMFiT Method for German Language☆15May 10, 2019Updated 7 years ago
- 🤖 AI GitHub App that automatically reviews PRs, triages issues, and monitors repository health using LLMs.☆23Updated this week
- 🦛 CHONK docs with Chonkie ✨ — The lightweight ingestion library for fast, efficient and robust RAG pipelines☆4,656Updated this week
- One library to split them all: Sentence, Code, Docs. Chunk smarter, not harder — built for LLMs, RAG pipelines, and beyond.☆84Updated this week
- MODE (Mixture of Document Experts) is an advanced RAG framework that enhances query response quality by combining hierarchical document c…☆16Sep 3, 2025Updated 11 months ago
- An opinionated list of practical tools for Conceptual Modeling and Linked Data☆40Mar 24, 2026Updated 4 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- GyShell is An Strong AI Agent Powered Terminal, Support SSH connection.☆42Updated this week
- pdfLLM is a completely open source, proof of concept RAG app.☆188Sep 1, 2025Updated 11 months ago
- Hybrid RAG system combining vector search, knowledge graph (LightRAG), and cross-encoder reranking — with Docling document parsing, visua…☆421Apr 20, 2026Updated 3 months ago
- Retrieval-augmented generation (RAG) for remote & local LLM use☆45May 24, 2025Updated last year
- [H] HyperspaceDB is a high-performance, vector database. It features 1-bit quantization, async replication, and native support for hierar…☆148Jul 27, 2026Updated 2 weeks ago
- Geometric AI research: a proven cube-math core, reusable vector-collapse dynamics, and reproducible experiments in embeddings, NLI, gener…☆15Jul 31, 2026Updated last week
- A Claude skill that builds PPTX decks matching your real designer decks: extracted design DNA, measured word budget, render-QA loop.☆28Jul 11, 2026Updated last month