Official PyTorch implementation and models for paper "Diffusion Beats Autoregressive in Data-Constrained Settings". We find diffusion models are significantly more data-efficient than standard left to right autoregressive models, due to their ability to learn from different token orderings.
☆128Jan 10, 2026Updated 8 months ago
Alternatives and similar repositories for diffusion-data-constraint
Users that are interested in diffusion-data-constraint are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The official github repo for "Diffusion Language Models are Super Data Learners".☆228Nov 6, 2025Updated 10 months ago
- The open-source materials for paper "Sparsing Law: Towards Large Language Models with Greater Activation Sparsity".☆32Nov 12, 2024Updated last year
- [NeurIPS 2024] Simple and Effective Masked Diffusion Language Model☆715Sep 29, 2025Updated 11 months ago
- [ICLR 2025] Official Pytorch Implementation of "Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN" by Pengxia…☆30Jul 24, 2025Updated last year
- Awesome Visual Tokenizers/Autoencoders☆20Nov 19, 2025Updated 10 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- UniDisc: A discrete diffusion model for joint multimodal generation, enabling controllable and efficient text-image synthesis, editing, a…☆143Apr 2, 2025Updated last year
- Official PyTorch implementation for ICLR2025 paper "Scaling up Masked Diffusion Models on Text"☆387Dec 22, 2024Updated last year
- Official PyTorch Implementation of "Latent Denoising Makes Good Visual Tokenizers"☆200Feb 24, 2026Updated 6 months ago
- Landing repository for the paper "Predicting the Order of Upcoming Tokens Improves Language Modeling"☆48May 13, 2026Updated 4 months ago
- Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture. Training an MDM using GPT with this repo!☆37Jun 23, 2025Updated last year
- ☆184Jan 8, 2026Updated 8 months ago
- Simple & Scalable Pretraining for Neural Architecture Research☆347Mar 31, 2026Updated 5 months ago
- [ICLR 2026] Official repository of "Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language Models"☆175Feb 16, 2026Updated 7 months ago
- [ICML 2026 Spotlight] On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models☆167Jun 8, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Towards Holistic evaluation of Generative Diffusion Transformers!☆105Jul 1, 2026Updated 2 months ago
- Code for paper: "Learning Diffusion Models with Flexible Representation Guidance"☆17Mar 18, 2026Updated 6 months ago
- [COLM 2025] Official code for "When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoni…☆15Oct 31, 2025Updated 10 months ago
- RND1: Scaling Diffusion Language Models☆188Aug 17, 2026Updated last month
- [COLM 2025: 1st Workshop on the Application of LLM Explainability to Reasoning and Planning] Latent Chain-of-Thought? Decoding the Depth-…☆22Aug 19, 2026Updated last month
- FlexTok: Resampling Images into 1D Token Sequences of Flexible Length☆333Sep 11, 2026Updated last week
- [ICLR 2025 Oral] Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models☆1,033Jul 10, 2025Updated last year
- Diffusion Language Models For Code Infilling Beyond Fixed-size Canvas☆119Feb 3, 2026Updated 7 months ago
- Code for ICML 2025 Paper "Highly Compressed Tokenizer Can Generate Without Training"☆206Jun 10, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Code for the paper Don't Pay Attention☆60Sep 25, 2025Updated 11 months ago
- Dream 7B, a large diffusion language model☆1,270Nov 21, 2025Updated 9 months ago
- Official Implementation of LaViDa: :A Large Diffusion Language Model for Multimodal Understanding☆230Dec 17, 2025Updated 9 months ago
- UQ: Assessing Language Models on Unsolved Questions☆30Aug 26, 2025Updated last year
- Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders☆263Feb 13, 2026Updated 7 months ago
- ☆31Aug 18, 2025Updated last year
- ☆54May 20, 2024Updated 2 years ago
- [ICLR2025] DiffuGPT and DiffuLLaMA: Scaling Diffusion Language Models via Adaptation from Autoregressive Models☆404May 31, 2025Updated last year
- ☆32Sep 4, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- MMaDA - Open-Sourced Multimodal Large Diffusion Language Models (dLLMs with block diffusion, mixed-CoT, unified RL)☆1,671Feb 14, 2026Updated 7 months ago
- Learn from Your Mistakes: Self-Correcting Masked Diffusion Models☆16Jun 25, 2026Updated 2 months ago
- FROM $f(x)$ AND $g(x)$ TO $f(g(x))$: LLMs Learn New Skills in RL by Composing Old Ones☆72Jan 26, 2026Updated 7 months ago
- [ICLR 2026] Discrete Diffusion Forcing (D2F): dLLMs Can Do Faster-Than-AR Inference☆262Feb 3, 2026Updated 7 months ago
- Measuring the Signal to Noise Ratio in Language Model Evaluation☆32Aug 19, 2025Updated last year
- [ICML 2026] Code for Equilibrium Reasoners: learning attractor dynamics for scalable reasoning☆52Updated this week
- See https://github.com/cuda-mode/triton-index/ instead!☆11May 8, 2024Updated 2 years ago