High performance inference engine for diffusion models
☆107Sep 5, 2025Updated 10 months ago
Alternatives and similar repositories for DAX
Users that are interested in DAX are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Distributed parallel 3D-Causal-VAE for efficient training and inference☆50Aug 20, 2025Updated 11 months ago
- Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation☆133Apr 28, 2026Updated 2 months ago
- ☆32Jul 2, 2025Updated last year
- ☆52May 19, 2025Updated last year
- https://wavespeed.ai/ Context parallel attention that accelerates DiT model inference with dynamic caching☆427Jul 5, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- 📚A curated list of Awesome Diffusion Inference Papers with Codes: Sampling, Cache, Quantization, Parallelism, etc.🎉☆578Jun 13, 2026Updated last month
- A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.☆1,234Updated this week
- NVSHMEM‑Tutorial: Build a DeepEP‑like GPU Buffer☆195Feb 11, 2026Updated 5 months ago
- KsanaDiT: High-Performance DiT (Diffusion Transformer) Inference Framework for Video & Image Generation☆62May 13, 2026Updated 2 months ago
- NCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.☆23Apr 10, 2026Updated 3 months ago
- xDiT: A Scalable Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism☆2,659Jul 14, 2026Updated last week
- ☆24Sep 4, 2025Updated 10 months ago
- High Performance LLM Inference Operator Library☆1,041Updated this week
- FlashTile is a CUDA Tile IR compiler that is compatible with NVIDIA's tileiras, targeting SM70 through SM121 NVIDIA GPUs.☆61Feb 6, 2026Updated 5 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆16Sep 12, 2023Updated 2 years ago
- An experimental communicating attention kernel based on DeepEP.☆34Jul 29, 2025Updated 11 months ago
- Some funny cute/cuteDSL code snippets☆33Mar 2, 2026Updated 4 months ago
- [ICML 2025] XAttention: Block Sparse Attention with Antidiagonal Scoring☆280Jul 6, 2025Updated last year
- FastCache: Fast Caching for Diffusion Transformer Through Learnable Linear Approximation [Efficient ML Model]☆52Apr 29, 2026Updated 2 months ago
- ☆148Aug 18, 2025Updated 11 months ago
- [HPCA 2026] A GPU-optimized system for efficient long-context LLMs decoding with low-bit KV cache.☆96May 14, 2026Updated 2 months ago
- Quantized Attention on GPU☆45Nov 22, 2024Updated last year
- [ASPLOS'26] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter☆174Feb 27, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Low overhead tracing library and trace visualizer for pipelined CUDA kernels☆137Updated this week
- 使用 cutlass 仓库在 ada 架构上实现 fp8 的 flash attention☆82Aug 12, 2024Updated last year
- A Triton JIT runtime and ffi provider in C++☆37Updated this week
- Wan2.2-Lightning: Speed up wan2.2 model with distillation☆304Nov 7, 2025Updated 8 months ago
- ☆45Oct 15, 2025Updated 9 months ago
- A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training☆883Updated this week
- Autonomous GPU kernel optimization system driven by AI agents.☆31Mar 29, 2026Updated 3 months ago
- Cute layout visualization☆43Jan 18, 2026Updated 6 months ago
- DeepSeek-V3.2-Exp DSA Warmup Lightning Indexer training operator based on tilelang☆47Nov 19, 2025Updated 8 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Getting Started with Triton: A Tutorial for Python Beginners☆61Mar 26, 2026Updated 3 months ago
- Distributed Compiler based on Triton for Parallel Systems☆1,494Updated this week
- [ICCV 2025] The official implementation of "Neighboring Autoregressive Modeling for Efficient Visual Generation"☆62Apr 5, 2025Updated last year
- [ICLR 2026] Official implementation of DiCache: Let Diffusion Model Determine Its Own Cache☆61Jan 26, 2026Updated 5 months ago
- A high-throughput and memory-efficient inference and serving engine for LLMs☆17Jun 3, 2024Updated 2 years ago
- [ICCV2025] From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers☆406Mar 2, 2026Updated 4 months ago
- A unified inference and post-training framework for accelerated video generation.☆3,862Updated this week