Organize the Web: Constructing Domains Enhances Pre-Training Data Curation
☆83May 2, 2025Updated last year
Alternatives and similar repositories for WebOrganizer
Users that are interested in WebOrganizer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆41May 26, 2026Updated last month
- ☆20Jun 27, 2026Updated 3 weeks ago
- Official repository for MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models [NeurIPS 2024]☆80Nov 14, 2024Updated last year
- Aioli: A unified optimization framework for language model data mixing☆33Jan 17, 2025Updated last year
- [ICML 2025] Predictive Data Selection: The Data That Predicts Is the Data That Teaches☆66Mar 4, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICML 2024] Selecting High-Quality Data for Training Language Models☆204Dec 8, 2025Updated 7 months ago
- [ICLR 2025] 🧬 RegMix: Data Mixture as Regression for Language Model Pre-training (Spotlight)☆194Feb 17, 2025Updated last year
- Data mapping framework for rust stuff☆56Mar 25, 2026Updated 3 months ago
- Official Code Repository for [AutoScale📈: Scale-Aware Data Mixing for Pre-Training LLMs] Published as a conference paper at **COLM 2025*…☆14Aug 8, 2025Updated 11 months ago
- Code for ICML 25 paper "Metadata Conditioning Accelerates Language Model Pre-training (MeCo)"☆51Jun 30, 2025Updated last year
- Reproducible, flexible LLM evaluations☆388Mar 24, 2026Updated 3 months ago
- Official Repository of Paper "Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs"☆15Sep 25, 2025Updated 9 months ago
- Revisiting Mid-training in the Era of Reinforcement Learning Scaling☆189Jul 23, 2025Updated 11 months ago
- DataComp for Language Models☆1,454Sep 9, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The official repository for SkyLadder: Better and Faster Pretraining via Context Window Scheduling☆43Dec 29, 2025Updated 6 months ago
- Language models scale reliably with over-training and on downstream tasks☆102Apr 2, 2024Updated 2 years ago
- Debiasing Through Data Attribution☆13May 23, 2024Updated 2 years ago
- ☆39Apr 17, 2024Updated 2 years ago
- CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generation☆14Aug 19, 2025Updated 11 months ago
- ☆28Mar 21, 2024Updated 2 years ago
- ☆53Jan 24, 2024Updated 2 years ago
- ☆23Jun 30, 2026Updated 3 weeks ago
- What's In My Big Data (WIMBD) - a toolkit for analyzing large text datasets☆229Nov 16, 2024Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Experimental tl;dr summaries for datasets on the Hugging Face Hub!☆10Apr 4, 2024Updated 2 years ago
- Ongoing research project for code&math LLMs☆32Jul 4, 2025Updated last year
- The simplest, fastest repository for training/finetuning medium-sized GPTs.☆199Jan 19, 2026Updated 6 months ago
- Implementation of stop sequencer for Huggingface Transformers☆16Jun 6, 2023Updated 3 years ago
- A Survey on Data Selection for Language Models☆260Apr 29, 2025Updated last year
- An automated data pipeline scaling RL to pretraining levels☆76Jun 2, 2026Updated last month
- DSIR large-scale data selection framework for language model training☆275Apr 7, 2024Updated 2 years ago
- ☆20Feb 18, 2025Updated last year
- Tooling for exact and MinHash deduplication of large-scale text datasets☆90Mar 24, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICLR 2026] Official Implementation of ProxyThinker: Test-Time Guidance through Small Visual Reasoners.☆22Sep 24, 2025Updated 9 months ago
- ☆21Jun 27, 2024Updated 2 years ago
- ☆44Oct 13, 2023Updated 2 years ago
- ☆52Mar 9, 2026Updated 4 months ago
- [NeurIPS-2023] The PyTorch Implementation of MoSo. The algorithms are based on our paper: "Data Pruning via Moving-one-Sample-out". MoSo …☆10May 21, 2026Updated 2 months ago
- Code and training scripts for FlexOlmo☆151Apr 20, 2026Updated 3 months ago
- Pytorch implementation of DoReMi, a method for optimizing the data mixture weights in language modeling datasets☆357Dec 26, 2023Updated 2 years ago