Repo for training MLMs, CLMs, or T5-type models on the OLM pretraining data, but it should work with any hugging face text dataset.
☆98Feb 9, 2023Updated 3 years ago
Alternatives and similar repositories for olm-training
Users that are interested in olm-training are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Pipeline for pulling and processing online language model pretraining data from the web☆179Jul 31, 2023Updated 3 years ago
- LTG-Bert☆35Jan 8, 2024Updated 2 years ago
- A tiny BERT for low-resource monolingual models☆32Dec 24, 2025Updated 8 months ago
- Experiments for XLM-V Transformers Integeration☆13Feb 8, 2023Updated 3 years ago
- Repo for ICML23 "Why do Nearest Neighbor Language Models Work?"☆59Jan 12, 2023Updated 3 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- **ARCHIVED** Filesystem interface to 🤗 Hub☆60Apr 6, 2023Updated 3 years ago
- [ICML 2023] Exploring the Benefits of Training Expert Language Models over Instruction Tuning☆98Apr 26, 2023Updated 3 years ago
- Experiments for efforts to train a new and improved t5☆76Apr 15, 2024Updated 2 years ago
- ☆16Mar 3, 2024Updated 2 years ago
- An open collection of implementation tips, tricks and resources for training large language models☆506Mar 8, 2023Updated 3 years ago
- Plug-and-play Search Interfaces with Pyserini and Hugging Face☆31Aug 5, 2023Updated 3 years ago
- Code repo for "Model-Generated Pretraining Signals Improves Zero-Shot Generalization of Text-to-Text Transformers" (ACL 2023)☆22Nov 1, 2023Updated 2 years ago
- Local emulator for Hugging Face Inference Endpoints customer handlers☆26Updated this week
- Training and evaluation code for the paper "Headless Language Models: Learning without Predicting with Contrastive Weight Tying" (https:/…☆30Apr 17, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆30Sep 27, 2021Updated 4 years ago
- A Streamlit app to add structured tags to a dataset card☆23Jun 30, 2022Updated 4 years ago
- Code Roberta version of RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder☆10Mar 16, 2023Updated 3 years ago
- [ACL'24 Oral] Analysing The Impact of Sequence Composition on Language Model Pre-Training☆24Aug 18, 2024Updated 2 years ago
- Binary Passage Retriever (BPR) - an efficient passage retriever for open-domain question answering☆176Jun 6, 2021Updated 5 years ago
- ☆13Mar 27, 2020Updated 6 years ago
- Long-context pretrained encoder-decoder models☆97Oct 28, 2022Updated 3 years ago
- Official code and model checkpoints for our EMNLP 2022 paper "RankGen - Improving Text Generation with Large Ranking Models" (https://arx…☆140Aug 2, 2023Updated 3 years ago
- A repository to get acquainted with basic training tasks in natural language processing and machine learning☆11Dec 27, 2023Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆102Dec 17, 2022Updated 3 years ago
- A package for fine tuning of pretrained NLP transformers using Semi Supervised Learning☆14Oct 27, 2021Updated 4 years ago
- Staged Training for Transformer Language Models☆33Mar 31, 2022Updated 4 years ago
- Hugging Face and Pyserini interoperability☆20May 18, 2023Updated 3 years ago
- Adding new tasks to T0 without catastrophic forgetting☆33Oct 20, 2022Updated 3 years ago
- Fast & Simple repository for pre-training and fine-tuning T5-style models☆1,021Aug 21, 2024Updated 2 years ago
- This repository contains the code for paper Prompting ELECTRA Few-Shot Learning with Discriminative Pre-Trained Models.☆48Jun 7, 2022Updated 4 years ago
- Official repository for "Scaling Retrieval-Based Langauge Models with a Trillion-Token Datastore".☆226Dec 16, 2025Updated 9 months ago
- Enhaced version of Wikiextrator: A wikipedia dumps extractor☆31Sep 17, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code for WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models.☆93Sep 12, 2024Updated 2 years ago
- PyTorch + HuggingFace code for RetoMaton: "Neuro-Symbolic Language Modeling with Automaton-augmented Retrieval" (ICML 2022), including an…☆288Oct 20, 2022Updated 3 years ago
- Efficient few-shot learning with Sentence Transformers☆2,813Updated this week
- Scaling Data-Constrained Language Models☆345Jun 28, 2025Updated last year
- ☆50Mar 14, 2024Updated 2 years ago
- Codes and files for the paper Are Emergent Abilities in Large Language Models just In-Context Learning☆33Jan 9, 2025Updated last year
- ☆16Jan 12, 2023Updated 3 years ago