Source code to accompany research paper on training multi token prediction language models using self-distillation.
☆42Feb 21, 2026Updated 7 months ago
Alternatives and similar repositories for mtp-lm
Users that are interested in mtp-lm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for "DynaGuard: A Dynamic Guardrail Model With User-Defined Policies."☆25Nov 3, 2025Updated 10 months ago
- (ECCV 2026): Official code for Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models☆23Jul 9, 2026Updated 2 months ago
- ☆20Nov 4, 2025Updated 10 months ago
- Code for Dayal Kalra's research internship on scalable curvature measures for neural networks.☆29Feb 3, 2026Updated 7 months ago
- Code for the paper Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs☆70Jun 23, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Official PyTorch Implementation for Learning a Generative Meta-Model of LLM Activations, ICML 2026☆95Sep 12, 2026Updated last week
- Official implementation of "Learning To Draft: Adaptive Speculative Decoding with Reinforcement Learning" (ICLR 2026)☆23Mar 1, 2026Updated 6 months ago
- PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation (ICLR 26)☆35Jun 10, 2026Updated 3 months ago
- ☆24Jun 16, 2026Updated 3 months ago
- [ICML 2026] Hybrid Policy Distillation (HPD) is a practical distillation framework for reasoning-oriented language models. This repositor…☆25Apr 24, 2026Updated 4 months ago
- 📄 Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay☆25Jul 17, 2026Updated 2 months ago
- Gemstones: A Model Suite for Multi-Faceted Scaling Laws (NeurIPS 2025)☆35Sep 28, 2025Updated 11 months ago
- SCT: An Efficient Self-Supervised Cross-View Training For Sentence Embedding (TACL)☆16Jul 27, 2024Updated 2 years ago
- ☆66Jul 3, 2026Updated 2 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [ICML2026] Reproduce Kimi K1.5/K2 RL algorithm and rollout system☆21Apr 9, 2026Updated 5 months ago
- Cross-GPU KV Cache Marketplace☆27Nov 12, 2025Updated 10 months ago
- Official Implementation of MARS☆30Apr 21, 2026Updated 5 months ago
- ☆20Sep 6, 2025Updated last year
- [ICLR 2026] Official code for BézierFlow: Learning Bézier Stochastic Interpolant Schedulers for Few-Step Generation☆24Apr 13, 2026Updated 5 months ago
- A selective knowledge distillation algorithm for efficient speculative decoders☆39Nov 27, 2025Updated 9 months ago
- A Teacher–Student Cooperation Framework to Synthesize Student-Consistent SFT Data☆37May 1, 2026Updated 4 months ago
- On Policy Distillation Build on top of Verl☆100Sep 3, 2026Updated 2 weeks ago
- Accelerating RL for LLM Reasoning with Optimal Advantage Regression☆41May 30, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- PyTorch Implementation of Image Generation with a Sphere Encoder☆47May 20, 2026Updated 4 months ago
- PyTorch implementation for "Generative Modeling on Manifolds Through Mixture of Riemannian Diffusion Processes" (ICML 2024).☆13Jul 21, 2024Updated 2 years ago
- On the Quartic Invariant of Odd Degree Binary Forms — paper, Lean formalization, and computational verification☆16Apr 16, 2026Updated 5 months ago
- Auditing agents for fine-tuning safety☆22Oct 21, 2025Updated 11 months ago
- Code for the paper "Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns"☆19Mar 15, 2024Updated 2 years ago
- ProAct is a framework designed to enable Large Language Model (LLM) agents to perform accurate, multi-turn lookahead reasoning in interac…☆18Feb 11, 2026Updated 7 months ago
- Inverse Scaling in Test-Time Compute☆26Dec 3, 2025Updated 9 months ago
- ☆47Jan 30, 2026Updated 7 months ago
- [EMNLP 2026] MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Scaling of Diffusion Language Models☆30Sep 11, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- An MLX implementation of Meta AI's ESM-2 protein language model☆16Aug 16, 2025Updated last year
- FS-DFM: Fast and Accurate Long Text Generation with Few-Step Diffusion Language Models. FS-DFM accepted for ICLR 2026☆50Sep 11, 2026Updated last week
- ☆19May 25, 2026Updated 3 months ago
- [ICML 2026] Official implementation for paper: Learning Self-Correction in Vision–Language Models via Rollout Augmentation☆16Jun 4, 2026Updated 3 months ago
- The official implementation of NOSA☆19Jun 11, 2026Updated 3 months ago
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 6 months ago
- ☆42May 26, 2026Updated 3 months ago