Simple-to-use scoring function for arbitrarily tokenized texts.
☆51Feb 19, 2025Updated last year
Alternatives and similar repositories for tokenization-scorer
Users that are interested in tokenization-scorer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This is the repository for MorphScore, a tokenizer evaluation framework for morphological alignment.☆17Jul 10, 2025Updated last year
- Code and data for "Superbizarre Is Not Superb: Derivational Morphology Improves BERT's Interpretation of Complex Words"☆18Aug 17, 2021Updated 4 years ago
- ☆11Mar 17, 2026Updated 4 months ago
- TokEval: intrinsic quality metrics for tokenizers across natural language, code, and math☆46Jul 4, 2026Updated 3 weeks ago
- [NeurIPS 2022] Your Transformer May Not be as Powerful as You Expect (official implementation)☆35Aug 6, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Easy-to-use framework for evaluating cross-lingual consistency of factual knowledge (Supported LLaMA, BLOOM, mT5, RoBERTa, etc.) Paper he…☆28Aug 8, 2025Updated 11 months ago
- GC4LM: A Colossal (Biased) language model for German☆13May 2, 2021Updated 5 years ago
- Cross-Linguistic Norms, Ratings, and Relations for Words and Concepts☆16Updated this week
- Code used to create the Linked WikiText-2 dataset☆16May 22, 2023Updated 3 years ago
- ☆17Jun 9, 2025Updated last year
- ☆47Feb 5, 2023Updated 3 years ago
- This repository contains code for the paper "Better Estimation of the KL Divergence Between Language Models"☆19May 30, 2025Updated last year
- [ACL 2025] 🔍 Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment☆11Apr 6, 2025Updated last year
- ☆79Apr 29, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Code for the papers "Induction of Subgoal Automata for Reinforcement Learning" (AAAI-20) and "Induction and Exploitation of Subgoal Autom…☆14Aug 15, 2023Updated 2 years ago
- Data and code: "Answering legal questions from laymen in German civil law system", Büttner & Habernal, EACL'24☆17Mar 2, 2024Updated 2 years ago
- ☆10May 14, 2024Updated 2 years ago
- Compositional Obverter Communication Learning From Raw Visual Input - Pytorch Implementation☆19Jul 8, 2021Updated 5 years ago
- A curated collection of resources for prompt engineering, optimization, and automatic prompt generation across text, image, video, and mu…☆18Sep 24, 2025Updated 10 months ago
- [ICML 2026] Improving GPT via a simple normalization strategy☆15May 22, 2026Updated 2 months ago
- Learning to Annotate Part Segmentation with Gradient Matching (ICLR 2022)☆12Apr 26, 2022Updated 4 years ago
- Minimal code to train ELMo models in recent versions of TensorFlow☆14Jun 16, 2026Updated last month
- Repository for the WACV 2024 paper "PsyMo: A Dataset for Estimating Self-Reported Psychological Traits from Gait"☆14Feb 22, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A software for transferring pre-trained English models to foreign languages☆20Mar 20, 2023Updated 3 years ago
- Resources related to EMNLP 2021 paper "FAME: Feature-Based Adversarial Meta-Embeddings for Robust Input Representations"☆13Dec 14, 2021Updated 4 years ago
- A collection of resources for PII detection, anonymization, privacy-preserving techniques, and GDPR compliance in Large Language Model (L…☆19Sep 24, 2025Updated 10 months ago
- ACL 2021 paper "Style is NOT a single variable: Case Studies for Cross-Style Language Understanding " by Dongyeop Kang and Eduard Hovy☆15Jul 19, 2021Updated 5 years ago
- ALBERT trained on Mongolian text corpus☆19Jan 10, 2021Updated 5 years ago
- Official repository for the paper "Approximating Two-Layer Feedforward Networks for Efficient Transformers"☆39Jun 11, 2025Updated last year
- ✱ Understanding the underlying learning dynamics of simple tasks in Transformer networks☆19Aug 16, 2024Updated last year
- Code for EMNLP2021 paper "Allocating Large Vocabulary Capacity for Cross-lingual Language Model Pre-training"☆20Nov 12, 2021Updated 4 years ago
- Efficient Language Model Training through Cross-Lingual and Progressive Transfer Learning☆30Jan 25, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Hosting the JSON for the GPT4 Tokenizer☆63Apr 6, 2023Updated 3 years ago
- [NAACL 2024] A Framework aims to wisely initialize unseen subword embeddings in PLMs for efficient large-scale continued pretraining☆18Nov 26, 2023Updated 2 years ago
- Dialogue Act classification☆18Jan 15, 2024Updated 2 years ago
- CC: Causality-Aware Coverage Criterion for Deep Neural Networks☆12Feb 15, 2023Updated 3 years ago
- Analyse des Pegida facebook Korpus☆10Jan 31, 2015Updated 11 years ago
- This is a german ELMo deep contextualized word representation. It is trained on a special German Wikipedia Text Corpus.☆28Dec 15, 2019Updated 6 years ago
- ☆53Jan 18, 2024Updated 2 years ago