π₯ Fast State-of-the-Art Tokenizers optimized for Research and Production
β11,158Oct 7, 2026Updated this week
Alternatives and similar repositories for tokenizers
Users that are interested in tokenizers are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π€ The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation toolsβ22,042Updated this week
- π€ Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal modelβ¦β167,023Updated this week
- Unsupervised text tokenizer for Neural Network-based text generation.β12,118Updated this week
- π A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (iβ¦β9,910Updated this week
- Minimalist ML framework for Rustβ21,150Updated this week
- GPUs on demand by Runpod - Special Offer Available β’ AdRun AI, ML, and HPC workloads on powerful cloud GPUsβwithout limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Rust native ready-to-use NLP pipelines and transformer-based models (BERT, DistilBERT, GPT2,...)β3,077Jan 13, 2026Updated 8 months ago
- Facebook AI Research Sequence-to-Sequence Toolkit written in Python.β32,213Sep 30, 2025Updated last year
- Papers & presentation materials from Hugging Face's internal science dayβ2,052Oct 31, 2020Updated 5 years ago
- State-of-the-Art Embeddings, Retrieval, and Rerankingβ19,158Updated this week
- A very simple framework for state-of-the-art Natural Language Processing (NLP)β14,391Oct 27, 2025Updated 11 months ago
- An open-source NLP research library, built on PyTorch.β11,885Nov 22, 2022Updated 3 years ago
- Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.β31,390Updated this week
- A library for efficient similarity search and clustering of dense vectors.β41,107Updated this week
- Code for the paper "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer"β6,556Updated this week
- End-to-end encrypted email - Proton Mail β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- π Accelerate inference and training of π€ Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimizationβ¦β3,498Updated this week
- Rust bindings for the C++ api of PyTorch.β5,498Aug 23, 2026Updated last month
- Simple, safe way to store and distribute tensorsβ3,913Updated this week
- DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.β43,208Updated this week
- Extremely fast Query Engine for DataFrames, written in Rustβ40,015Updated this week
- Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and moreβ36,387Updated this week
- π« Industrial-strength Natural Language Processing (NLP) in Pythonβ33,950Sep 30, 2026Updated last week
- Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.β16,035Updated this week
- πͺβKnock Knock: Get notified when your training ends with only two additional lines of codeβ2,826Jun 23, 2023Updated 3 years ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Qdrant - High-performance, massive-scale Vector Database and Vector Search Engine for the next generation of AI. Also available in the clβ¦β34,978Updated this week
- Repository to track the progress in Natural Language Processing (NLP), including the datasets and the current state-of-the-art for the moβ¦β22,948Jul 28, 2024Updated 2 years ago
- Tantivy is a full-text search engine library inspired by Apache Lucene and written in Rustβ16,192Updated this week
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generatorsβ2,365Mar 23, 2024Updated 2 years ago
- tiktoken is a fast BPE tokeniser for use with OpenAI's models.β19,413Aug 17, 2026Updated last month
- TensorFlow code and pre-trained models for BERTβ40,044Jul 23, 2024Updated 2 years ago
- π€ PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.β21,769Updated this week
- Stanford NLP Python library for tokenization, sentence segmentation, NER, and parsing of many human languagesβ7,894Updated this week
- Large Language Model Text Generation Inferenceβ10,880Mar 21, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [Unmaintained, see README] An ecosystem of Rust libraries for working with large language modelsβ6,155Jun 24, 2024Updated 2 years ago
- Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalitiesβ22,228Sep 21, 2026Updated 2 weeks ago
- Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.β43,992Updated this week
- Accessible large language models via k-bit quantization for PyTorch.β8,514Sep 7, 2026Updated last month
- Rust bindings for the Python interpreterβ16,207Updated this week
- A Rust machine learning framework.β4,755Aug 22, 2026Updated last month
- XLNet: Generalized Autoregressive Pretraining for Language Understandingβ6,188May 28, 2023Updated 3 years ago