⚡️Qwen-Image 4.8x🎉 speedup with Hybrid Acceleration for low VRAM GPUs
☆17Oct 24, 2025Updated 11 months ago
Alternatives and similar repositories for qwen-image-fast
Users that are interested in qwen-image-fast are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 🔥LongCat-Video 1.7x🎉 speedup: cache acceleration and 4/8-bits weight only.☆16Oct 28, 2025Updated 11 months ago
- High performance RMSNorm Implement by using SM Core Storage(Registers and Shared Memory)☆31Jan 22, 2026Updated 8 months ago
- VibeRL is a Reinforcement Learning framework built essentially through vibe coding with Kimi K2.☆18Sep 28, 2026Updated last week
- Xmixers: A collection of SOTA efficient token/channel mixers☆29Sep 4, 2025Updated last year
- A forked version of flux-fast that makes flux-fast even faster with cache-dit, 3.3x speedup on NVIDIA L20.☆24Jul 18, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- lightNet (Object Detection and Semantic Segmentation) for ONNX and TensorRT☆16Jul 4, 2023Updated 3 years ago
- Collection of Acceleration Methods for Generative AI☆29Dec 9, 2025Updated 10 months ago
- Muon in Int8 Precision Made Possible☆20Jun 18, 2026Updated 3 months ago
- A demonstrative example of running SGLang Diffusion with DP router☆26Mar 15, 2026Updated 6 months ago
- ☆17Nov 14, 2023Updated 2 years ago
- High-throughput tensor loading for PyTorch☆268Aug 17, 2026Updated last month
- HiCache: Hermite Polynomial-based Feature Cache for diffusion inference☆15Sep 7, 2026Updated last month
- Detection and Tracking ROS node based on CenterPoint and Kalman Filter☆24Feb 24, 2024Updated 2 years ago
- ☆21Oct 17, 2025Updated 11 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- GEMV implementation with CUTLASS☆21Aug 21, 2025Updated last year
- ☆11May 24, 2024Updated 2 years ago
- DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching☆25Apr 15, 2026Updated 5 months ago
- PyTorch implementation of the Flash Spectral Transform Unit.☆23Sep 19, 2024Updated 2 years ago
- c++实现的clip推理,模型有一点点改动,但是不大,改动和导出模型的代码可以在readme里找到,模型文件都在Releases里,包括AX650的模型。新增支持ChineseCLIP☆31Jun 19, 2025Updated last year
- SGEMM optimization with cuda step by step☆23Mar 23, 2024Updated 2 years ago
- A curated list of all awesome pygames created by Agneay B Nair☆12Apr 28, 2024Updated 2 years ago
- Agent-native Seedance 2.0 short-film studio: cli for AI, canvas for human☆16Sep 8, 2026Updated last month
- ☆23Nov 6, 2025Updated 11 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆35Jul 2, 2025Updated last year
- Qwen-Image's DiT inference with TensorRT-10☆21Oct 13, 2025Updated 11 months ago
- [MLSys 26] 🥇 Solution for Gated Delta Net Track of MLSys 26 Flash infer competition☆36May 22, 2026Updated 4 months ago
- fake CUTLASS to get peformance☆25Apr 28, 2026Updated 5 months ago
- [CVPR 2025] The official implementation of "CacheQuant: Comprehensively Accelerated Diffusion Models"☆51Nov 2, 2025Updated 11 months ago
- FlashTile is a CUDA Tile IR compiler that is compatible with NVIDIA's tileiras, targeting SM70 through SM121 NVIDIA GPUs.☆60Feb 6, 2026Updated 8 months ago
- FLA but cuTile☆27Apr 17, 2026Updated 5 months ago
- ☆23Aug 20, 2025Updated last year
- Kernel Library for Large Headdim Attention (64~1024, BF16/FP8/FP4), 1.5x~15x↑ vs PyTorch SDPA.☆336Sep 26, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆31Aug 25, 2023Updated 3 years ago
- ☆23Aug 14, 2024Updated 2 years ago
- a size profiler for cuda binary☆73Updated this week
- Event matching for log records☆11May 12, 2014Updated 12 years ago
- Autonomous GPU kernel optimization system driven by AI agents.☆32Mar 29, 2026Updated 6 months ago
- [WIP] Better (FP8) attention for Hopper☆33Aug 21, 2026Updated last month
- Minimal PyTorch implementation of TP, SP, FSDP and sharded-EMA☆32Nov 27, 2025Updated 10 months ago