Stop using static chunk sizes. A lightweight, production-ready RAG ingestion toolkit. Uses Docling for layout-aware parsing and applies smart heuristics for optimal chunking (PDF vs Code vs MD). Extracted from a production RAG platform
☆71Mar 15, 2026Updated 4 months ago
Alternatives and similar repositories for smart-ingest-kit
Users that are interested in smart-ingest-kit are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A simple CPU only OCR for pdf/images/word/excel to markdown. With streamlit.☆52Jan 26, 2026Updated 5 months ago
- A Docker-powered RAG system that understands the difference between code and prose. Ingest your codebase and documentation, then query th…☆351Mar 14, 2026Updated 4 months ago
- A Python CLI to test, benchmark, and find the best RAG chunking strategy for your Markdown documents.☆114Jan 18, 2026Updated 6 months ago
- An opinionated, "speed" and "usability" focused agentic TUI with a built-in MCP registry/plugin system.☆28Apr 14, 2026Updated 3 months ago
- Exploring retrieval systems for language models☆14Apr 12, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A simple streamlit app to play with qwen3-2b-VL to perform OCR. Dockerized set up, tested with 3060 12 GB.☆32Nov 23, 2025Updated 7 months ago
- Enterprise-grade Retrieval-Augmented Generation system with microservices architecture.☆23Mar 15, 2026Updated 4 months ago
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆17Mar 20, 2026Updated 4 months ago
- Token-aware, LangChain-compatible semantic chunker with PDF, markdown, and layout support☆13Jun 28, 2025Updated last year
- REFRAG: LLM-powered representations for better RAG retrieval. Improve precision, reduce context size, same speed.☆29Dec 29, 2025Updated 6 months ago
- Privacy-focused, self-hosted RAG assistant for querying codebases with local or cloud LLMs.☆21May 29, 2026Updated last month
- SmartRAG is a privacy-first multimodal RAG system that lets you chat intelligently with your documents, images, and audio. Upload PDFs, W…☆110Apr 6, 2026Updated 3 months ago
- GrantFlow.ai is a platform for creating grant applications using ML and AI☆58Mar 20, 2026Updated 4 months ago
- Professional RAG development skills for Claude Code - audit, evaluate, optimize, and scaffold RAG pipelines☆33Jan 18, 2026Updated 6 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆23Mar 21, 2026Updated 4 months ago
- Elixir library to generate Ecto migrations from a PostgreSQL schema SQL file. Uses NimbleParsec and macro-style code generation.☆18Dec 12, 2025Updated 7 months ago
- Laddr is a python framework for building multi-agent systems where agents communicate, delegate tasks, and execute work in parallel. Thin…☆342Jul 14, 2026Updated last week
- CustomGPT.ai’s RAG API’s Starter Kit, including multi-instance embedded widgets, floating buttons, and standalone application.☆52Dec 23, 2025Updated 6 months ago
- Multi-strategy RAG system achieving 74% Recall@10 on MultiHop-RAG. Combines RAPTOR hierarchical retrieval, knowledge graphs, HyDE, BM25, …☆41Feb 3, 2026Updated 5 months ago
- PipesHub is an open-source fully extensible AI context layer that unifies your business data for explainable enterprise search and agenti…☆3,042Updated this week
- ☆12Feb 23, 2024Updated 2 years ago
- ☆16Feb 3, 2026Updated 5 months ago
- A fully local, zero-API, zero-finetune multi-agent AI architecture that makes an 8B base model perform high level model reasoning, resear…☆35Nov 25, 2025Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- TreeThinkerAgent is a lightweight orchestration layer that turns any LLM into an autonomous multi-step reasoning agent. It supports multi…☆21Feb 11, 2026Updated 5 months ago
- ☆53Nov 21, 2025Updated 8 months ago
- WeMush Open Labeling Standard☆48Jan 30, 2026Updated 5 months ago
- A Paperless-ngx consume script that leverages Docling to provide superior OCR and layout analysis for PDFs, Office documents, and images.☆18Dec 7, 2025Updated 7 months ago
- ☆64Jul 9, 2026Updated last week
- Validated, private, shareable knowledge-graph memory for AI — per-tenant, write-gated, PostgreSQL-authoritative, served over MCP.☆18Updated this week
- This is more like a QA system☆16Updated this week
- CLI tool for intelligent Obsidian vault interaction using RAG. Query your notes with natural language and convert URLs to markdown files …☆36Oct 16, 2025Updated 9 months ago
- ☆27Jun 22, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- A Developmental AI That Learns Like a Child☆29Mar 29, 2026Updated 3 months ago
- Development tool for Model Context Protocol servers☆13Apr 23, 2025Updated last year
- Middleware for AI Agents that verifies grounding and prevents hallucinations. Returns structured retry suggestions for self-correction.☆51Dec 11, 2025Updated 7 months ago
- this is an easy way to make ai podcast useing ai loccaly like ollama and the tts of piper☆16Feb 17, 2026Updated 5 months ago
- Reduce your OpenClaw agent costs. Free real-time LLM cost tracking + dashboard. Installs in 60 seconds.☆15Mar 15, 2026Updated 4 months ago
- Give your local LLM a real memory with a lightweight, fully local memory system. 100% offline and under your control.☆77Sep 16, 2025Updated 10 months ago
- ☆75May 2, 2025Updated last year