Code for collecting, processing, and preparing datasets for the Common Pile
☆259Feb 11, 2026Updated 6 months ago
Alternatives and similar repositories for common-pile
Users that are interested in common-pile are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17Nov 26, 2024Updated last year
- [ICML'25] "Rethinking Addressing in Language Models via Contextualized Equivariant Positional Encoding" by Jiajun Zhu, Peihao Wang, Ruisi…☆15Jun 6, 2025Updated last year
- Few-shot Learning with Auxiliary Data☆31Dec 8, 2023Updated 2 years ago
- Code for the paper "BPE stays on SCRIPT", "Which Pieces Does Unigram Tokenization Really Need?" and MinGram☆22Aug 27, 2026Updated last week
- ☆12Jul 6, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆39Apr 17, 2024Updated 2 years ago
- [ACL 2025] 🔍 Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment☆11Apr 6, 2025Updated last year
- OLMost every training recipe you need to perform data interventions with the OLMo family of models.☆75Jul 21, 2026Updated last month
- TokEval: intrinsic quality metrics for tokenizers across natural language, code, and math☆53Aug 31, 2026Updated last week
- ☆79Dec 7, 2023Updated 2 years ago
- Pile Deduplication Code☆18May 15, 2023Updated 3 years ago
- Code for boomerang distillation enables zero-shot model size interpolation.☆22Jul 10, 2026Updated last month
- Luth is a state-of-the-art series of fine-tuned LLMs for French☆47Oct 12, 2025Updated 10 months ago
- The Institutional Data Initiative's pipeline for analyzing, refining, and publishing the Institutional Books datasets.☆57Aug 19, 2026Updated 2 weeks ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A polite and user-friendly downloader for Common Crawl data☆94Aug 5, 2026Updated last month
- QAmeleon introduces synthetic multilingual QA data using PaLM, a 540B large language model. This dataset was generated by prompt tuning P…☆34Aug 15, 2023Updated 3 years ago
- data related codebase for polyglot project☆19Mar 30, 2023Updated 3 years ago
- Dataset: BuzzFeed News “Trending” Strip, 2018–2023☆18May 24, 2023Updated 3 years ago
- Intersectional bias in hate speech and abusive language datasets☆15Jan 25, 2024Updated 2 years ago
- Official code for the paper CipherDAug: Ciphertext based Data Augmentation for Neural Machine Translation published at ACL 2022 main conf…☆12Apr 6, 2023Updated 3 years ago
- Collection of academic works in natural language processing, computational linguistics, and computational cognitive science that study th…☆23Mar 20, 2024Updated 2 years ago
- PathPiece tokenizer☆14Nov 10, 2024Updated last year
- Utilities for PyTorch distributed☆26Feb 27, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ☆16Aug 5, 2025Updated last year
- ☆16Feb 10, 2026Updated 6 months ago
- Code for NeurIPS 2024 Spotlight: "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations"☆93Oct 30, 2024Updated last year
- ☆263Oct 27, 2025Updated 10 months ago
- A lightweight, user-friendly data-plane for LLM training.☆41Sep 10, 2025Updated 11 months ago
- ☆311Apr 23, 2025Updated last year
- Gemini API for OCR☆15Nov 17, 2025Updated 9 months ago
- Data and tools for generating and inspecting OLMo pre-training data.☆1,541Aug 24, 2026Updated 2 weeks ago
- Repository of data on web domains.☆19May 24, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This project studies the performance and robustness of language models and task-adaptation methods.☆154May 18, 2024Updated 2 years ago
- A pipeline for LLM knowledge distillation☆117May 7, 2026Updated 4 months ago
- Legible, Scalable, Reproducible Foundation Models with Named Tensors and Jax☆714Jan 26, 2026Updated 7 months ago
- A tool for detecting viruses and NSFW material in WARC files☆18Jul 15, 2026Updated last month
- Adversarial Tokenization☆39Nov 21, 2025Updated 9 months ago
- A curated list of tools, guides and resources for the Replicate AI model platform☆17Jan 10, 2024Updated 2 years ago
- The pipeline for the OSCAR corpus☆178Nov 9, 2025Updated 9 months ago