π‘ 30x faster tokenization for every HuggingFace model
β55Aug 27, 2026Updated last month
Alternatives and similar repositories for tokie
Users that are interested in tokie are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π Want one client library for all your embeddings? π Choose Catsu! π±β83Apr 21, 2026Updated 5 months ago
- Nearly Inference Free Embeddings: make your RAG queries 500x fasterβ88Apr 27, 2026Updated 5 months ago
- π¦ Chonkie's recipes for Agents, Context Engineering, and more! π§βπ³ Chonkie knows how to cook (and it cooks well!)β18Feb 28, 2026Updated 7 months ago
- FlexiTokensβ23Dec 27, 2025Updated 9 months ago
- Orchestration layer to run multiple Copilot tasks concurrently with Temporal Workflowsβ18Feb 7, 2026Updated 7 months ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- benchmarks for LLM tokenizersβ21Mar 25, 2026Updated 6 months ago
- Next-generation Punkt sentence boundary detection with zero dependenciesβ32Updated this week
- Simple customizable evaluation for text retrieval performance of Sentence Transformers embedders on PDFsβ30Jan 20, 2025Updated last year
- Official Repository for "Hypencoder: Hypernetworks for Information Retrieval"β41Sep 20, 2025Updated last year
- Trainable embedding transformation for confidence estimation, feature extraction, explainability and conversion from dense to sparse.β28Jun 23, 2026Updated 3 months ago
- Knowledgeable Embedding: Injecting dynamically updatable entity knowledge into embeddings to enhance RAGβ15Aug 31, 2025Updated last year
- Make-like task executor for Unix OSβ17Jul 28, 2026Updated 2 months ago
- Model implementation for the contextual embeddings projectβ52Jun 2, 2025Updated last year
- High performance implementation of the WARP (SIGIR'25) retrieval engine.β36Sep 8, 2026Updated 3 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- TokEval: intrinsic quality metrics for tokenizers across natural language, code, and mathβ55Aug 31, 2026Updated last month
- bb25 is a fast, self-contained BM25 + Bayesian calibration implementation with a minimal Python API.β150Mar 17, 2026Updated 6 months ago
- Fetch all the docs you need for your vibe coding projectsβ23Aug 23, 2025Updated last year
- The fastest BM25 scoring engine: 2,300x faster than BM25S. 28K QPS on 8.8M docs. 5 BM25 variants (Robertson, Lucene, ATIRE, BM25L, BM25+β¦β51Apr 3, 2026Updated 6 months ago
- [ACL 20] Probing Linguistic Features of Sentence-level Representations in Neural Relation Extractionβ13Apr 21, 2020Updated 6 years ago
- Vocabulary Trimming (VT) is a model compression technique, which reduces a multilingual LM vocabulary to a target language by deleting irβ¦β70Oct 25, 2024Updated last year
- Efficient BM25 with DuckDB π¦β69Dec 20, 2024Updated last year
- Python Module implementing SRPβ12Jul 29, 2022Updated 4 years ago
- Send your invoices and expenses via email, let AI handle themβ24Nov 7, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- https://footprints.baulab.infoβ17Oct 4, 2024Updated last year
- Official repository of TACHIOM.β65Sep 20, 2026Updated last week
- This is a prototype of a multi-lingual suite for named-entity recognition in Python. β‘οΈ The project has moved to: https://gitlab.opencodeβ¦β21Mar 20, 2026Updated 6 months ago
- Sparse Embedding Compression for Scalable Retrieval in Recommender Systemsβ39Nov 21, 2025Updated 10 months ago
- [SIGIR 2025] The official repo for "Scaling Sparse and Dense Retrieval in Decoder-Only LLMs"β22Mar 31, 2025Updated last year
- Personal README for Github profile.β10Jun 12, 2023Updated 3 years ago
- Multi-Figurative Language Generation (COLING 2022)β12Jan 30, 2023Updated 3 years ago
- β13Nov 15, 2017Updated 8 years ago
- My NER Experiments with ModernBERT and Ettinβ30Jul 17, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A code snippet that proves that there is no legal, potentially non-reachable chess position with more than 218 moves.β23May 23, 2024Updated 2 years ago
- ANE accelerated embedding models!β20Dec 11, 2024Updated last year
- HSEB: Hybrid Search Engine Benchmarkβ21Oct 5, 2025Updated 11 months ago
- RelEx - A simple framework for Relation Extraction built on AllenNLPβ15Jun 17, 2020Updated 6 years ago
- π The Fastest Chunker in the West πΊπΈ Upto 1TB/s "semantic" chunking, quick and easy!β360May 28, 2026Updated 4 months ago
- Python implementation of Levenshtein distance and Levenshtein automata matchingβ27May 8, 2019Updated 7 years ago
- Learning BPE embeddings by first learning a segmentation model and then training word2vecβ19Dec 18, 2022Updated 3 years ago