☆85Apr 20, 2026Updated 4 months ago
Alternatives and similar repositories for dolma3
Users that are interested in dolma3 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Tooling for exact and MinHash deduplication of large-scale text datasets☆96Mar 24, 2026Updated 5 months ago
- Data mapping framework for rust stuff☆59Mar 25, 2026Updated 5 months ago
- PyTorch building blocks for the OLMo ecosystem☆1,519Updated this week
- Reproducible, flexible LLM evaluations☆394Mar 24, 2026Updated 5 months ago
- OLMost every training recipe you need to perform data interventions with the OLMo family of models.☆75Jul 21, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICML 2025] Predictive Data Selection: The Data That Predicts Is the Data That Teaches☆66Mar 4, 2025Updated last year
- An automated data pipeline scaling RL to pretraining levels☆76Jun 2, 2026Updated 3 months ago
- Organize the Web: Constructing Domains Enhances Pre-Training Data Curation☆83May 2, 2025Updated last year
- decontamination☆38Mar 4, 2026Updated 6 months ago
- ☆33Jul 22, 2025Updated last year
- Code for ICML 25 paper "Metadata Conditioning Accelerates Language Model Pre-training (MeCo)"☆53Jun 30, 2025Updated last year
- ☆33Oct 15, 2025Updated 10 months ago
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆35May 26, 2026Updated 3 months ago
- Learning in Noisy MDP (which is governed by stochastic, exogenous input processes) with input-dependent baseline☆10Aug 7, 2020Updated 6 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Codebase for the work “Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?”☆75Apr 14, 2026Updated 4 months ago
- A high-throughput and memory-efficient inference and serving engine for LLMs☆21Feb 1, 2026Updated 7 months ago
- SORTED: A curated collection of interesting ideas, tools, and resources in neuroscience, data management, and data science, all in the sp…☆27Aug 10, 2025Updated last year
- Experimentation with Streamlit for personal LLM tool☆15Jun 19, 2023Updated 3 years ago
- Data Efficacy for Language Model Training☆52May 29, 2026Updated 3 months ago
- Theorem relational dependencies automatic extraction and visualization as a graph for Lean4.☆54Feb 15, 2026Updated 6 months ago
- Open source implementation of DataRater (https://arxiv.org/abs/2505.17895)☆26Sep 20, 2025Updated 11 months ago
- ☆106Jul 16, 2026Updated last month
- Code and data for "An Accurate Unsupervised Method for Joint Entity Alignment and Dangling Entity Detection".☆15Mar 26, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML…☆208Mar 29, 2026Updated 5 months ago
- [AAAI 2025] Augmenting Math Word Problems via Iterative Question Composing (https://arxiv.org/abs/2401.09003)☆23Oct 2, 2025Updated 11 months ago
- Scaling is a distributed training library and installable dependency designed to scale up neural networks, with a dedicated module for tr…☆66Nov 18, 2025Updated 9 months ago
- [ACL 2026 Main] Official Repo for Paper "Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Ali…☆16Jul 1, 2026Updated 2 months ago
- [COLM 2025: 1st Workshop on the Application of LLM Explainability to Reasoning and Planning] Latent Chain-of-Thought? Decoding the Depth-…☆21Aug 19, 2026Updated 2 weeks ago
- ☆41May 26, 2026Updated 3 months ago
- Automatic textbook formalization of Grinberg Algebraic Combinatorics☆17Jul 28, 2026Updated last month
- Source codes for paper "BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity".☆19Jan 10, 2026Updated 7 months ago
- ☆19Sep 6, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Kinematic and dynamic models of continuum and articulated soft robots.☆16Aug 13, 2026Updated 3 weeks ago
- [EMNLP 2025🔥] UNComp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design from an Uncertainty-Aware Perspective☆20Jan 7, 2026Updated 8 months ago
- ☆16Aug 5, 2025Updated last year
- ☆23Jun 12, 2025Updated last year
- ☆21May 18, 2026Updated 3 months ago
- The official code of "Mano: Restriking Manifold Optimization for LLM Training".☆25Jun 1, 2026Updated 3 months ago
- AllenAI's post-training codebase☆3,858Updated this week