Organize the Web: Constructing Domains Enhances Pre-Training Data Curation
☆84May 2, 2025Updated last year
Alternatives and similar repositories for WebOrganizer
Users that are interested in WebOrganizer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆42May 26, 2026Updated 4 months ago
- [ICML 2025] Predictive Data Selection: The Data That Predicts Is the Data That Teaches☆66Mar 4, 2025Updated last year
- Aioli: A unified optimization framework for language model data mixing☆34Jan 17, 2025Updated last year
- ☆21Jun 27, 2026Updated 3 months ago
- Official repository for MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models [NeurIPS 2024]☆80Nov 14, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICML 2024] Selecting High-Quality Data for Training Language Models☆203Dec 8, 2025Updated 10 months ago
- [ICLR 2025] 🧬 RegMix: Data Mixture as Regression for Language Model Pre-training (Spotlight)☆215Feb 17, 2025Updated last year
- Code for ICML 25 paper "Metadata Conditioning Accelerates Language Model Pre-training (MeCo)"☆53Jun 30, 2025Updated last year
- Data mapping framework for rust stuff☆60Mar 25, 2026Updated 6 months ago
- Official Code Repository for [AutoScale📈: Scale-Aware Data Mixing for Pre-Training LLMs] Published as a conference paper at **COLM 2025*…☆14Aug 8, 2025Updated last year
- Reproducible, flexible LLM evaluations☆399Mar 24, 2026Updated 6 months ago
- Official Repository of Paper "Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs"☆15Sep 25, 2025Updated last year
- Revisiting Mid-training in the Era of Reinforcement Learning Scaling☆188Jul 23, 2025Updated last year
- DataComp for Language Models☆1,475Sep 9, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- The official repository for SkyLadder: Better and Faster Pretraining via Context Window Scheduling☆44Dec 29, 2025Updated 9 months ago
- Language models scale reliably with over-training and on downstream tasks☆104Apr 2, 2024Updated 2 years ago
- Debiasing Through Data Attribution☆13May 23, 2024Updated 2 years ago
- ☆39Apr 17, 2024Updated 2 years ago
- CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generation☆14Aug 19, 2025Updated last year
- ☆28Mar 21, 2024Updated 2 years ago
- ☆52Jan 24, 2024Updated 2 years ago
- ☆23Jun 30, 2026Updated 3 months ago
- Simple and scalable tools for data-driven pretraining data selection.☆31Jun 9, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Experimental tl;dr summaries for datasets on the Hugging Face Hub!☆10Apr 4, 2024Updated 2 years ago
- Ongoing research project for code&math LLMs☆32Jul 4, 2025Updated last year
- The simplest, fastest repository for training/finetuning medium-sized GPTs.☆204Jan 19, 2026Updated 8 months ago
- Implementation of stop sequencer for Huggingface Transformers☆16Jun 6, 2023Updated 3 years ago
- A Survey on Data Selection for Language Models☆261Apr 29, 2025Updated last year
- DSIR large-scale data selection framework for language model training☆277Apr 7, 2024Updated 2 years ago
- An automated data pipeline scaling RL to pretraining levels☆76Jun 2, 2026Updated 4 months ago
- ☆20Feb 18, 2025Updated last year
- Tooling for exact and MinHash deduplication of large-scale text datasets☆98Mar 24, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Optimizing Anytime Reasoning via Budget Relative Policy Optimization☆54Jul 15, 2025Updated last year
- [ICLR 2026] Official Implementation of ProxyThinker: Test-Time Guidance through Small Visual Reasoners.☆22Sep 24, 2025Updated last year
- ☆17Sep 11, 2026Updated 3 weeks ago
- [NeurIPS-2023] The PyTorch Implementation of MoSo. The algorithms are based on our paper: "Data Pruning via Moving-one-Sample-out". MoSo …☆10May 21, 2026Updated 4 months ago
- ☆52Mar 9, 2026Updated 7 months ago
- ☆57Apr 11, 2024Updated 2 years ago
- Pytorch implementation of DoReMi, a method for optimizing the data mixture weights in language modeling datasets☆360Dec 26, 2023Updated 2 years ago