Code for collecting, processing, and preparing datasets for the Common Pile
☆260Feb 11, 2026Updated 7 months ago
Alternatives and similar repositories for common-pile
Users that are interested in common-pile are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17Nov 26, 2024Updated last year
- [ICML'25] "Rethinking Addressing in Language Models via Contextualized Equivariant Positional Encoding" by Jiajun Zhu, Peihao Wang, Ruisi…☆15Jun 6, 2025Updated last year
- Few-shot Learning with Auxiliary Data☆31Dec 8, 2023Updated 2 years ago
- Selected code and data for The Online Books Page and related applications☆12Sep 1, 2026Updated 3 weeks ago
- Code for the paper "BPE stays on SCRIPT", "Which Pieces Does Unigram Tokenization Really Need?" and MinGram☆22Updated this week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- DKPro C4CorpusTools is a collection of tools for processing CommonCrawl corpus, including Creative Commons license detection, boilerplate…☆54Jun 12, 2020Updated 6 years ago
- ☆26Aug 19, 2025Updated last year
- ☆39Apr 17, 2024Updated 2 years ago
- [ACL 2025] 🔍 Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment☆11Apr 6, 2025Updated last year
- OLMost every training recipe you need to perform data interventions with the OLMo family of models.☆76Jul 21, 2026Updated 2 months ago
- ☆78Dec 7, 2023Updated 2 years ago
- Pile Deduplication Code☆18May 15, 2023Updated 3 years ago
- Luth is a state-of-the-art series of fine-tuned LLMs for French☆47Oct 12, 2025Updated 11 months ago
- The Institutional Data Initiative's pipeline for analyzing, refining, and publishing the Institutional Books datasets.☆60Aug 19, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A polite and user-friendly downloader for Common Crawl data☆96Aug 5, 2026Updated last month
- DPO, but faster 🚀☆51Dec 6, 2024Updated last year
- QAmeleon introduces synthetic multilingual QA data using PaLM, a 540B large language model. This dataset was generated by prompt tuning P…☆34Aug 15, 2023Updated 3 years ago
- Dataset: BuzzFeed News “Trending” Strip, 2018–2023☆18May 24, 2023Updated 3 years ago
- Intersectional bias in hate speech and abusive language datasets☆15Jan 25, 2024Updated 2 years ago
- Collection of academic works in natural language processing, computational linguistics, and computational cognitive science that study th…☆24Mar 20, 2024Updated 2 years ago
- Using experimental methods to merge large language models☆11Jul 11, 2026Updated 2 months ago
- PathPiece tokenizer☆14Nov 10, 2024Updated last year
- ☆33Jul 8, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆16Feb 10, 2026Updated 7 months ago
- ☆17Aug 23, 2025Updated last year
- Code for NeurIPS 2024 Spotlight: "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations"☆94Oct 30, 2024Updated last year
- Highly concurrent and fast content processing for Mighty Inference Server☆10Feb 6, 2023Updated 3 years ago
- ☆265Oct 27, 2025Updated 11 months ago
- A lightweight, user-friendly data-plane for LLM training.☆42Sep 10, 2025Updated last year
- Tooling for exact and MinHash deduplication of large-scale text datasets☆97Mar 24, 2026Updated 6 months ago
- SKT A.X LLM 3.1☆13Jul 24, 2025Updated last year
- This project studies the performance and robustness of language models and task-adaptation methods.☆154May 18, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆22Apr 27, 2026Updated 5 months ago
- Legible, Scalable, Reproducible Foundation Models with Named Tensors and Jax☆716Jan 26, 2026Updated 8 months ago
- Python binding for gumbo-parser using Cython☆14Aug 16, 2016Updated 10 years ago
- A pipeline for LLM knowledge distillation☆119May 7, 2026Updated 4 months ago
- A tool for detecting viruses and NSFW material in WARC files☆18Jul 15, 2026Updated 2 months ago
- Modded vLLM to run pipeline parallelism over public networks☆41May 20, 2025Updated last year
- [EMNLP 2023] 💬 Language Identification with Support for More Than 2000 Labels☆219Apr 15, 2026Updated 5 months ago