Code for the paper "Getting the most out of your tokenizer for pre-training and domain adaptation"
☆22Feb 14, 2024Updated 2 years ago
Alternatives and similar repositories for tokenizer-bench
Users that are interested in tokenizer-bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 0-Shot Tokenizer Transplant☆14May 16, 2025Updated last year
- Multi-Word Probabilistic based supertokenizer☆15May 15, 2025Updated last year
- FactNews is the first dataset to predict sentence-level factuality of news reporting. Furthemore, we provide baseline results for sentenc…☆12Jun 12, 2025Updated last year
- Master thesis: Exploring bias in German NLG (GPT-3 & GerPT-2). Applies regard classification and bias mitigation triggers.☆16Sep 25, 2024Updated last year
- Temporary remove unused tokens during training to save ram and speed.☆23Jun 15, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization [ACL 2026]☆20Apr 18, 2026Updated 4 months ago
- Ukrainian ELECTRA model☆12Mar 11, 2023Updated 3 years ago
- ☆16May 14, 2024Updated 2 years ago
- [ACL 2025] 🔍 Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment☆11Apr 6, 2025Updated last year
- Efficient Language Model Training through Cross-Lingual and Progressive Transfer Learning☆30Jan 25, 2023Updated 3 years ago
- ☆10Nov 8, 2023Updated 2 years ago
- A scalable benchmark for state representation learning in visual reinforcement learning.☆17Jun 23, 2025Updated last year
- Code for SaGe subword tokenizer (EACL 2023)☆28Nov 30, 2024Updated last year
- [ICML 2026] Improving GPT via a simple normalization strategy☆15May 22, 2026Updated 3 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Minimal code to train ELMo models in recent versions of TensorFlow☆14Jun 16, 2026Updated 2 months ago
- Resources related to EMNLP 2021 paper "FAME: Feature-Based Adversarial Meta-Embeddings for Robust Input Representations"☆13Dec 14, 2021Updated 4 years ago
- Statistics on multilingual datasets☆17Jul 12, 2022Updated 4 years ago
- Implementation of the paper 'Sentence Bottleneck Autoencoders from Transformer Language Models'☆17Mar 14, 2022Updated 4 years ago
- [NeurIPS 2025] MergeBench: A Benchmark for Merging Domain-Specialized LLMs☆50Aug 28, 2026Updated last week
- A byte-level decoder architecture that matches the performance of tokenized Transformers.☆68Apr 24, 2024Updated 2 years ago
- From Hero to Zéroe: A Benchmark of Low-Level Adversarial Attacks☆15Feb 23, 2023Updated 3 years ago
- ✂️ Sentence segmentation with wtpsplit's state-of-the-art Segment any Text (SaT) models☆41May 2, 2026Updated 4 months ago
- Python module to remove wiki markup text.☆10Jan 15, 2016Updated 10 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Code for "Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations" [NAACL Findings 2024]☆14Apr 3, 2026Updated 5 months ago
- Set-Equivariant Deep Learning Models☆22Dec 23, 2021Updated 4 years ago
- Code for the paper "Greed is All You Need: An Evaluation of Tokenizer Inference Methods"☆15Nov 26, 2024Updated last year
- Finite-state script normalization and processing utilities☆52Updated this week
- This repository provides the code for applying Contrastive Learning Penalty Loss (CLPL) and Mixture of Experts (MoE) to the BGE-M3 text e…☆11Dec 27, 2024Updated last year
- [ICLR 2025] MiniPLM: Knowledge Distillation for Pre-Training Language Models☆79Nov 23, 2024Updated last year
- Official repository for the paper "Approximating Two-Layer Feedforward Networks for Efficient Transformers"☆39Jun 11, 2025Updated last year
- ☆20Mar 30, 2022Updated 4 years ago
- A tool for extracting plain text from Wikipedia dumps☆15Sep 13, 2018Updated 7 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Typst linguistic examples with minimalist syntax☆18Updated this week
- Example for a Monty-enabled RLM in DSPy☆20Feb 16, 2026Updated 6 months ago
- Overview of corpora/datasets for Germanic low-resource languages and dialects. Accompanies "A Survey of Corpora for Germanic Low-Resource…☆28Feb 16, 2026Updated 6 months ago
- Micro-framework for publishing linked data☆11Aug 1, 2017Updated 9 years ago
- Find informative examples to efficiently (human)-evaluate NLG models.☆17Apr 22, 2026Updated 4 months ago
- Repository for "Self-Distillation for Model Stacking Unlocks Cross-Lingual NLU in 200+ Languages"☆15Oct 4, 2024Updated last year
- Implementation of Nested Named Entity Recognition using Flair☆24Oct 29, 2021Updated 4 years ago