Code implementation of synthetic continued pretraining
☆162Jan 6, 2025Updated last year
Alternatives and similar repositories for Synthetic_Continued_Pretraining
Users that are interested in Synthetic_Continued_Pretraining are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆574Nov 20, 2024Updated last year
- Code for Paper (Preserving Diversity in Supervised Fine-tuning of Large Language Models)☆59May 12, 2025Updated last year
- ☆15Oct 26, 2021Updated 4 years ago
- This repository presents the original implementation of Pretraining Data Detection for Large Language Models: A Divergence-based Calibrat…☆23May 21, 2025Updated last year
- A live reading list for LLM data synthesis (Updated to July, 2025).☆494Apr 9, 2026Updated 4 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [ICML 2024] Selecting High-Quality Data for Training Language Models☆203Dec 8, 2025Updated 8 months ago
- Predicting Out-of-Distribution Error with the Projection Norm☆19Jul 27, 2022Updated 4 years ago
- [ACL 2025] We introduce ScaleQuest, a scalable, novel and cost-effective data synthesis method to unleash the reasoning capability of LLM…☆69Oct 27, 2024Updated last year
- Generates and optimizes Haiku system and user prompts for classification☆15Oct 27, 2025Updated 10 months ago
- A reading list on LLM based Synthetic Data Generation 🔥☆1,551Jun 5, 2025Updated last year
- Official implementation of "Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data" (ICLR 2024)☆37Oct 16, 2024Updated last year
- [ICLR'25] DataGen: Unified Synthetic Dataset Generation via Large Language Models☆69Mar 8, 2025Updated last year
- Deita: Data-Efficient Instruction Tuning for Alignment [ICLR2024]☆602Dec 9, 2024Updated last year
- Data for "Datamodels: Predicting Predictions with Training Data"☆97May 25, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆15May 28, 2024Updated 2 years ago
- Code for the paper "The Journey, Not the Destination: How Data Guides Diffusion Models"☆26Dec 12, 2023Updated 2 years ago
- Official implementation of the paper "From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large L…☆55Jun 24, 2024Updated 2 years ago
- DSIR large-scale data selection framework for language model training☆276Apr 7, 2024Updated 2 years ago
- [COLM '24] Source-Aware Training Enables Knowledge Attribution in Language Models☆21Apr 1, 2025Updated last year
- Code for the paper "Data Feedback Loops: Model-driven Amplification of Dataset Biases"☆18Sep 9, 2022Updated 3 years ago
- ☆120Feb 18, 2025Updated last year
- MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer (EMNLP 2025)☆12Apr 18, 2025Updated last year
- ☆18Mar 2, 2026Updated 6 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆14Jun 13, 2025Updated last year
- DataComp for Language Models☆1,471Sep 9, 2025Updated 11 months ago
- ☆34Jul 11, 2024Updated 2 years ago
- Official implementation of SIGIR 2022 Paper "Task-Oriented Dialogue System as Natural Language Generation".☆14Apr 6, 2022Updated 4 years ago
- Revisiting Mid-training in the Era of Reinforcement Learning Scaling☆188Jul 23, 2025Updated last year
- Lightweight Adapting for Black-Box Large Language Models☆26Feb 15, 2024Updated 2 years ago
- InsTag: A Tool for Data Analysis in LLM Supervised Fine-tuning☆289Aug 20, 2023Updated 3 years ago
- ☆989Feb 7, 2025Updated last year
- [NeurIPS-2023] The PyTorch Implementation of MoSo. The algorithms are based on our paper: "Data Pruning via Moving-one-Sample-out". MoSo …☆10May 21, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Github repo for Peifeng's internship project☆13Nov 7, 2023Updated 2 years ago
- [NeurIPS 2023] Official Pytorch code for LOVM: Language-Only Vision Model Selection☆21Feb 3, 2024Updated 2 years ago
- Ongoing research project for code&math LLMs☆31Jul 4, 2025Updated last year
- ☆15Mar 4, 2022Updated 4 years ago
- Official Code Repository for [AutoScale📈: Scale-Aware Data Mixing for Pre-Training LLMs] Published as a conference paper at **COLM 2025*…☆14Aug 8, 2025Updated last year
- Repo for Rho-1: Token-level Data Selection & Selective Pretraining of LLMs.☆472Apr 18, 2024Updated 2 years ago
- Code for the paper "Quantifying Privacy Leakage in Graph Embedding" published in MobiQuitous 2020☆18Nov 11, 2021Updated 4 years ago