Source code to accompany research paper on training multi token prediction language models using self-distillation.
☆39Feb 21, 2026Updated 5 months ago
Alternatives and similar repositories for mtp-lm
Users that are interested in mtp-lm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆14Mar 2, 2025Updated last year
- Code for "DynaGuard: A Dynamic Guardrail Model With User-Defined Policies."☆23Nov 3, 2025Updated 8 months ago
- ☆19Nov 4, 2025Updated 8 months ago
- Code for Dayal Kalra's research internship on scalable curvature measures for neural networks.☆29Feb 3, 2026Updated 5 months ago
- Code for the paper Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs☆63Jun 23, 2026Updated last month
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆31May 21, 2026Updated 2 months ago
- Official PyTorch Implementation for Learning a Generative Meta-Model of LLM Activations, ICML 2026☆90Apr 30, 2026Updated 2 months ago
- Official implementation of "Learning To Draft: Adaptive Speculative Decoding with Reinforcement Learning" (ICLR 2026)☆19Mar 1, 2026Updated 4 months ago
- ☆20Jun 16, 2026Updated last month
- [ICML 2026] Hybrid Policy Distillation (HPD) is a practical distillation framework for reasoning-oriented language models. This repositor…☆24Apr 24, 2026Updated 2 months ago
- 📄 Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay☆24Updated this week
- Gemstones: A Model Suite for Multi-Faceted Scaling Laws (NeurIPS 2025)☆35Sep 28, 2025Updated 9 months ago
- ☆62Jul 3, 2026Updated 2 weeks ago
- PyTorch Implementation of Zero-Shot Vision Encoder Grafting via LLM Surrogates [ICCV'25]☆54Jul 10, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ICML2026] Reproduce Kimi K1.5/K2 RL algorithm and rollout system☆19Apr 9, 2026Updated 3 months ago
- Cross-GPU KV Cache Marketplace☆26Nov 12, 2025Updated 8 months ago
- 📄Small Batch Size Training for Language Models☆82Mar 18, 2026Updated 4 months ago
- [ICLR 2026] Official code for BézierFlow: Learning Bézier Stochastic Interpolant Schedulers for Few-Step Generation☆21Apr 13, 2026Updated 3 months ago
- A Teacher–Student Cooperation Framework to Synthesize Student-Consistent SFT Data☆33May 1, 2026Updated 2 months ago
- A selective knowledge distillation algorithm for efficient speculative decoders☆39Nov 27, 2025Updated 7 months ago
- [NeurIPS 2024] Goldfish Loss: Mitigating Memorization in Generative LLMs☆98Nov 17, 2024Updated last year
- ☆21Feb 10, 2025Updated last year
- On Policy Distillation Build on top of Verl☆92May 25, 2026Updated last month
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- PyTorch Implementation of Image Generation with a Sphere Encoder☆44May 20, 2026Updated 2 months ago
- PyTorch implementation for "Generative Modeling on Manifolds Through Mixture of Riemannian Diffusion Processes" (ICML 2024).☆13Jul 21, 2024Updated 2 years ago
- Auditing agents for fine-tuning safety☆21Oct 21, 2025Updated 9 months ago
- On the Quartic Invariant of Odd Degree Binary Forms — paper, Lean formalization, and computational verification☆16Apr 16, 2026Updated 3 months ago
- [NeurIPS 2024] Official implementation of the paper "Enhancing LLM’s Cognition via Structurization"☆24Aug 5, 2025Updated 11 months ago
- ProAct is a framework designed to enable Large Language Model (LLM) agents to perform accurate, multi-turn lookahead reasoning in interac…☆18Feb 11, 2026Updated 5 months ago
- Inverse Scaling in Test-Time Compute☆26Dec 3, 2025Updated 7 months ago
- ☆45Jan 30, 2026Updated 5 months ago
- MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Scaling of Diffusion Language Models☆27May 23, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- An MLX implementation of Meta AI's ESM-2 protein language model☆16Aug 16, 2025Updated 11 months ago
- FS-DFM: Fast and Accurate Long Text Generation with Few-Step Diffusion Language Models. FS-DFM accepted for ICLR 2026☆45Jan 6, 2026Updated 6 months ago
- ☆16May 25, 2026Updated last month
- The official implementation of NOSA☆19Jun 11, 2026Updated last month
- The official implementation of the paper "A Dual-Space Framework for General Knowledge Distillation of Large Language Models".☆18Jan 4, 2026Updated 6 months ago
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 4 months ago
- This repository contains code for the paper "Better Estimation of the KL Divergence Between Language Models"☆19May 30, 2025Updated last year