Stop using static chunk sizes. A lightweight, production-ready RAG ingestion toolkit. Uses Docling for layout-aware parsing and applies smart heuristics for optimal chunking (PDF vs Code vs MD). Extracted from a production RAG platform
☆71Mar 15, 2026Updated 4 months ago
Alternatives and similar repositories for smart-ingest-kit
Users that are interested in smart-ingest-kit are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- RAG for financial Analysis using SQL and calculator tool.☆16Apr 29, 2026Updated 2 months ago
- A simple CPU only OCR for pdf/images/word/excel to markdown. With streamlit.☆52Jan 26, 2026Updated 5 months ago
- A Docker-powered RAG system that understands the difference between code and prose. Ingest your codebase and documentation, then query th…☆350Mar 14, 2026Updated 4 months ago
- A Python CLI to test, benchmark, and find the best RAG chunking strategy for your Markdown documents.☆114Jan 18, 2026Updated 6 months ago
- Exploring retrieval systems for language models☆14Apr 12, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- pdfLLM is a completely open source, proof of concept RAG app.☆186Sep 1, 2025Updated 10 months ago
- A simple streamlit app to play with qwen3-2b-VL to perform OCR. Dockerized set up, tested with 3060 12 GB.☆32Nov 23, 2025Updated 7 months ago
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆17Mar 20, 2026Updated 4 months ago
- SmartRAG is a privacy-first multimodal RAG system that lets you chat intelligently with your documents, images, and audio. Upload PDFs, W…☆110Apr 6, 2026Updated 3 months ago
- REFRAG: LLM-powered representations for better RAG retrieval. Improve precision, reduce context size, same speed.☆29Dec 29, 2025Updated 6 months ago
- CustomGPT.ai’s RAG API’s Starter Kit, including multi-instance embedded widgets, floating buttons, and standalone application.☆52Dec 23, 2025Updated 6 months ago
- Privacy-focused, self-hosted RAG assistant for querying codebases with local or cloud LLMs.☆21May 29, 2026Updated last month
- ☆23Apr 3, 2026Updated 3 months ago
- GrantFlow.ai is a platform for creating grant applications using ML and AI☆58Mar 20, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- XLSX parser for LLMs, RAG, LangChain, LangGraph, CrewAI, Claude, MCP — turns Excel (.xlsx) into citation-ready JSON with formulas, charts…☆46Updated this week
- Professional RAG development skills for Claude Code - audit, evaluate, optimize, and scaffold RAG pipelines☆33Jan 18, 2026Updated 6 months ago
- An opinionated development framework for building production-ready AI agents with LangGraph. It grounds AI coding assistants (Cursor, Win…☆23May 20, 2026Updated 2 months ago
- ☆23Mar 21, 2026Updated 4 months ago
- Smart code bundler that turns repositories into optimized code bundles meeting a token budget in milliseconds☆42Feb 3, 2026Updated 5 months ago
- Save context for what matters. The last agent you'll ever need.☆33Sep 25, 2025Updated 9 months ago
- PipesHub is an open-source fully extensible AI context layer that unifies your business data for explainable enterprise search and agenti…☆3,038Updated this week
- Multi-strategy RAG system achieving 74% Recall@10 on MultiHop-RAG. Combines RAPTOR hierarchical retrieval, knowledge graphs, HyDE, BM25, …☆41Feb 3, 2026Updated 5 months ago
- ☆12Feb 23, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- TreeThinkerAgent is a lightweight orchestration layer that turns any LLM into an autonomous multi-step reasoning agent. It supports multi…☆21Feb 11, 2026Updated 5 months ago
- A fully local, zero-API, zero-finetune multi-agent AI architecture that makes an 8B base model perform high level model reasoning, resear…☆35Nov 25, 2025Updated 7 months ago
- ☆53Nov 21, 2025Updated 7 months ago
- Automate indexing files and pages from SharePoint into Azure AI Search☆11Jun 6, 2024Updated 2 years ago
- MCP server that allows Claude to have a voice.☆13May 5, 2025Updated last year
- ☆64Jul 9, 2026Updated last week
- Production-grade API to give your AI agents long-term memory without the boilerplate.☆72Dec 17, 2025Updated 7 months ago
- Validated, private, shareable knowledge-graph memory for AI — per-tenant, write-gated, PostgreSQL-authoritative, served over MCP.☆18Updated this week
- Cut LLM costs by up to 80% and unlock sub-millisecond responses with intelligent semantic caching.A drop-in, provider-agnostic LLM proxy …☆241Apr 24, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- This is more like a QA system☆16Jul 7, 2026Updated 2 weeks ago
- CLI tool for intelligent Obsidian vault interaction using RAG. Query your notes with natural language and convert URLs to markdown files …☆36Oct 16, 2025Updated 9 months ago
- ☆27Jun 22, 2025Updated last year
- Opinionated agentic RAG powered by LanceDB, Pydantic AI, and Docling☆547Updated this week
- Development tool for Model Context Protocol servers☆13Apr 23, 2025Updated last year
- ☆22May 17, 2026Updated 2 months ago
- Coordinate skills between Codex, Copilot, and Claude Code. Validates, analyzes, and syncs skills, subagents, commands, and configuration …☆68Jun 18, 2026Updated last month