Calculate perplexity on a text with pre-trained language models. Support MLM (eg. DeBERTa), recurrent LM (eg. GPT3), and encoder-decoder LM (eg. Flan-T5).
☆168Jun 20, 2025Updated last year
Alternatives and similar repositories for lmppl
Users that are interested in lmppl are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Analyzing mBERT's multilinguality in a small laboratory setting☆13Jun 12, 2023Updated 3 years ago
- 日本語文法誤り訂正ツール☆29Jun 22, 2022Updated 4 years ago
- ☆30Mar 20, 2024Updated 2 years ago
- Code Roberta version of RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder☆10Mar 16, 2023Updated 3 years ago
- Lite Self-Training☆30Jul 25, 2023Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Official code repository for Correct-N-Contrast☆22Jul 18, 2022Updated 4 years ago
- Word acquisition in neural language models (TACL 2022).☆21Jan 30, 2025Updated last year
- [NeurIPS 2022] Non-Linguistic Supervision for Contrastive Learning of Sentence Embeddings☆22Jan 30, 2023Updated 3 years ago
- ☆13Dec 1, 2021Updated 4 years ago
- Difference-based Contrastive Learning for Korean Sentence Embeddings☆23Mar 11, 2026Updated 4 months ago
- R library for accessing data from everypolitician.org☆20Apr 24, 2018Updated 8 years ago
- Official codebase for the NeurIPS 2023 paper: Towards Last-layer Retraining for Group Robustness with Fewer Annotations. https://arxiv.or…☆12May 15, 2024Updated 2 years ago
- ☆24Apr 8, 2019Updated 7 years ago
- A repository for experiments in quality-aware decoding☆18Jun 7, 2022Updated 4 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Benchmarking Large Language Models☆106Jun 20, 2025Updated last year
- ☆24Nov 22, 2022Updated 3 years ago
- albumentations test☆11Jun 23, 2020Updated 6 years ago
- This repository provides the code and dataset for the work published in the paper - Modeling Label Semantics for Predicting Emotional Rea…☆26Nov 8, 2020Updated 5 years ago
- A simple semi-supervised approach for creating huggingface data script loaders and upload to the hub.☆11Jun 23, 2024Updated 2 years ago
- Gated Pretrained Transformer model for robust denoised sequence-to-sequence modelling☆10May 29, 2021Updated 5 years ago
- Forked repo from https://github.com/EleutherAI/lm-evaluation-harness/commit/1f66adc☆81Feb 28, 2024Updated 2 years ago
- Pytorch Tutorial for M1 students. This repository include Encoder Deocder model and Classification model building code.☆12Jun 1, 2022Updated 4 years ago
- Code for EMNLP 2021 paper: Improving Sequence-to-Sequence Pre-training via Sequence Span Rewriting☆17Nov 30, 2021Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- HPLT Analytics☆15Jul 22, 2026Updated last week
- A library for evaluation of Grammatical Error Correction (GEC). Accepted to ACL'25 Demo: "gec-metrics: A Unified Library for Grammatical …☆14Jan 25, 2026Updated 6 months ago
- A accurate multilingual word aligner based on LaBSE☆24Oct 25, 2023Updated 2 years ago
- Convert English alphabet to Katakana☆15Jul 1, 2026Updated 3 weeks ago
- End-to-end codebase for finetuning LLMs (LLaMA 2, 3, etc.) with or without DP☆17Sep 23, 2024Updated last year
- ☆14Apr 8, 2018Updated 8 years ago
- Aspect based sentiment analysis for Hindi☆11Aug 31, 2017Updated 8 years ago
- Language Modelling Makes Sense - WSD (and more) with Contextual Embeddings☆97Jun 12, 2023Updated 3 years ago
- Tokenizer POS-Tagger and Dependency-parser with BERT/RoBERTa/DeBERTa/GPT models for Japanese and other languages☆55Feb 28, 2026Updated 5 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Python package implementing the greedy string tiling algorithm for comparing string similarity☆12Mar 20, 2023Updated 3 years ago
- ☆23Feb 26, 2024Updated 2 years ago
- ☆16Jul 26, 2024Updated 2 years ago
- Official repo of Progressive Data Expansion: data, code and evaluation☆29Nov 16, 2023Updated 2 years ago
- Code for the paper "Simple, Interpretable and Stable Method for Detecting Words with Usage Change across Corpora", ACL 2020.☆18Aug 28, 2020Updated 5 years ago
- CCL2022 领域问答库构建测评☆20Oct 31, 2022Updated 3 years ago
- Convenient Text-to-Text Training for Transformers☆18Dec 10, 2021Updated 4 years ago