This package, developed as part of our research detailed in the Chroma Technical Report, provides tools for text chunking and evaluation. It allows users to compare different chunking methods and includes implementations of several novel chunking strategies.
☆504Dec 13, 2025Updated 7 months ago
Alternatives and similar repositories for chunking_evaluation
Users that are interested in chunking_evaluation are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆34Jun 17, 2024Updated 2 years ago
- Efficient document processing for RAG using MaxMin Chunking.☆16May 20, 2026Updated 2 months ago
- Code for explaining and evaluating late chunking (chunked pooling)☆534Dec 23, 2024Updated last year
- Fast BM25 search in Python, powered by Numpy and Numba☆1,761Jul 22, 2026Updated 2 weeks ago
- ☆52Nov 18, 2025Updated 8 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A lightweight, low-dependency, unified API to use all common reranking and cross-encoder models.☆1,628Dec 20, 2025Updated 7 months ago
- ☆1,478Jun 18, 2024Updated 2 years ago
- Fast Multimodal Semantic Deduplication & Filtering☆958May 24, 2026Updated 2 months ago
- Easily use and train state of the art late-interaction retrieval methods (ColBERT) in any RAG pipeline. Designed for modularity and ease-…☆3,949May 17, 2025Updated last year
- ☆21Nov 26, 2024Updated last year
- Lite weight wrapper for the independent implementation of SPLADE++ models for search & retrieval pipelines. Models and Library created by…☆35Aug 24, 2024Updated last year
- Supercharge Your LLM Application Evaluations 🚀☆15,259Feb 24, 2026Updated 5 months ago
- ☆22Oct 14, 2024Updated last year
- 🦛 CHONK docs with Chonkie ✨ — The lightweight ingestion library for fast, efficient and robust RAG pipelines☆4,656Updated this week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- The Batched API provides a flexible and efficient way to process multiple requests in a batch, with a primary focus on dynamic batching o…☆161Jul 14, 2025Updated last year
- This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed not…☆29,010Jul 31, 2026Updated last week
- Late Interaction Models Training & Retrieval☆877Jul 23, 2026Updated 2 weeks ago
- UIC-2024☆12Apr 6, 2025Updated last year
- Model implementation for the contextual embeddings project☆48Jun 2, 2025Updated last year
- utilities for batched llm calls with retries☆51Jul 25, 2026Updated 2 weeks ago
- ☆12Feb 23, 2024Updated 2 years ago
- Use late-interaction multi-modal models such as ColPali in just a few lines of code.☆850Jan 28, 2025Updated last year
- A Python wrapper around HuggingFace's TGI (text-generation-inference) and TEI (text-embedding-inference) servers.☆32Sep 19, 2025Updated 10 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Starbucks: Improved Training for 2D Matryoshka Embeddings☆25Jun 30, 2025Updated last year
- ☆16Jun 26, 2026Updated last month
- Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean…☆15,291Updated this week
- Chunk your text using gpt4o-mini more accurately☆44Aug 3, 2024Updated 2 years ago
- Sandboxed tools and JS runtime for AI agents☆18Jul 13, 2026Updated 3 weeks ago
- Fast, Accurate, Lightweight Python library to make State of the Art Embedding☆3,134Updated this week
- Meta-Chunking: Learning Efficient Text Segmentation via Logical Perception☆277Sep 25, 2025Updated 10 months ago
- Overview and entry point for methods and experiment environments from paper "Stronger Baselines for Retrieval-Augmented Generation with L…☆25Nov 8, 2025Updated 9 months ago
- structured outputs for llms☆13,711Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆257Jun 10, 2025Updated last year
- Toolkit to segment text into sentences or other semantic units in a robust, efficient and adaptable way.☆1,327Updated this week
- Library for evaluating RAG using Nuclia's models☆19Jul 31, 2024Updated 2 years ago
- ☆197May 5, 2024Updated 2 years ago
- A toolkit to create optimal Production-readyRetrieval Augmented Generation(RAG) setup for your data☆1,541May 20, 2025Updated last year
- The LLM Evaluation Framework☆17,503Updated this week
- Get your documents ready for gen AI☆64,454Updated this week