☆31Aug 18, 2025Updated last year
Alternatives and similar repositories for LLADA_pretraining
Users that are interested in LLADA_pretraining are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The official github repo for "Training Optimal Large Diffusion Language Models", the first-ever large-scale diffusion language models sca…☆46Nov 6, 2025Updated 9 months ago
- SDAR (Synergy of Diffusion and AutoRegression), a large diffusion language model(1.7B, 4B, 8B, 30B)☆527Jul 29, 2026Updated last month
- LLaDA is a diffusion model for natural language that, unlike traditional autoregressive models, learns to model the distribution of text …☆28Jun 21, 2026Updated 2 months ago
- A Collection of Papers on Diffusion Language Models☆185Aug 2, 2026Updated last month
- [ACL '26] Source code for paper "Empirical Analysis of Decoding Biases in Masked Diffusion Models"☆45Jun 26, 2026Updated 2 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [ICLR 2026] Official code for TraceRL: Revolutionizing post-training for Diffusion LLMs, powering the SOTA TraDo series.☆520Jan 28, 2026Updated 7 months ago
- Official Implementation of wd1☆33Sep 25, 2025Updated 11 months ago
- [ICLR 2026] Discrete Diffusion Forcing (D2F): dLLMs Can Do Faster-Than-AR Inference☆261Feb 3, 2026Updated 7 months ago
- Open diffusion language model for code generation — releasing pretraining, evaluation, inference, and checkpoints.☆650Jul 20, 2026Updated last month
- A training-free framework that accelerates diffusion language models through dynamic cache eviction and sparse attention☆29Oct 16, 2025Updated 10 months ago
- GPU-optimized framework for training diffusion language models at any scale. The backend of Quokka, Super Data Learners, and OpenMoE 2 tr…☆344Nov 11, 2025Updated 9 months ago
- Methods and code for extending the context length of diffusion language models☆56Dec 7, 2025Updated 8 months ago
- Official PyTorch implementation of the paper "dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching" (dLLM-Cache…☆214May 1, 2026Updated 4 months ago
- Code for NeurIPS 2024 Paper "Fight Back Against Jailbreaking via Prompt Adversarial Tuning"☆22May 6, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [ICLR26] Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs☆33Dec 9, 2025Updated 8 months ago
- LightningRL: Breaking the Accuracy–Parallelism Trade-off of Block-wise dLLMs via Reinforcement Learning☆33Apr 25, 2026Updated 4 months ago
- See https://github.com/cuda-mode/triton-index/ instead!☆11May 8, 2024Updated 2 years ago
- Official Implementation of LaViDa: :A Large Diffusion Language Model for Multimodal Understanding☆229Dec 17, 2025Updated 8 months ago
- MDPO: Overcoming the Training-Inference Divide of Masked Diffusion Language Models☆46Jan 28, 2026Updated 7 months ago
- Automatically Update LLM Papers Daily using Github Actions. Ref: https://github.com/Vincentqyw/cv-arxiv-daily☆10Updated this week
- ☆110Nov 17, 2025Updated 9 months ago
- Official PyTorch implementation for ICLR2025 paper "Scaling up Masked Diffusion Models on Text"☆385Dec 22, 2024Updated last year
- Remasking Discrete Diffusion Models with Inference-Time Scaling☆78Feb 7, 2026Updated 6 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ASE2024] Mutual Learning-Based Framework for Enhancing Robustness of Code Models via Adversarial Training☆11Sep 13, 2024Updated last year
- ☆14Jun 24, 2024Updated 2 years ago
- On the Robustness of GUI Grounding Models Against Image Attacks☆12Apr 8, 2025Updated last year
- (ICLR 2026 🔥) Code for "The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs"☆79Feb 9, 2026Updated 6 months ago
- ☆21Oct 17, 2025Updated 10 months ago
- Dream 7B, a large diffusion language model☆1,266Nov 21, 2025Updated 9 months ago
- Official PyTorch implementation and models for paper "Diffusion Beats Autoregressive in Data-Constrained Settings". We find diffusion mod…☆128Jan 10, 2026Updated 7 months ago
- Official inference implementation of the paper "DON'T SETTLE TOO EARLY: SELF-REFLECTIVE REMASKING FOR DIFFUSION LANGUAGE MODELS". [ICLR 2…☆15Jan 28, 2026Updated 7 months ago
- Implementation of the paper "Improving the Accuracy-Robustness Trade-off of Classifiers via Adaptive Smoothing".☆10Feb 6, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [NeurIPS 2024]Repos for "Visualization-of-Thought" dataset, construction code and evaluation.☆37Oct 23, 2024Updated last year
- 🤗 [ICLR 2024] Disentangling Time Series Representations via Contrastive based l-Variational Inference☆20Dec 11, 2025Updated 8 months ago
- Tools to convert sigsep mus dataset from STEMS <-> WAV☆12Jul 15, 2020Updated 6 years ago
- [NeurIPS'25] dKV-Cache: The Cache for Diffusion Language Models☆136May 22, 2025Updated last year
- Official implemention for Diffusion Models Are Innate One-Step Generators☆27Jun 25, 2025Updated last year
- This is the official implementation for paper "On Powerful Ways to Generate: Autoregression, Diffusion, and Beyond".☆23Nov 17, 2025Updated 9 months ago
- Official Code for "Rethinking Diffusion Model in High Dimension"☆26May 20, 2025Updated last year