a simple Flash Attention v2 implementation with ROCM (RDNA3 GPU, roc wmma), mainly used for stable diffusion(ComfyUI) in Windows ZLUDA environments.
☆52Aug 25, 2024Updated last year
Alternatives and similar repositories for flash-attention-v2-RDNA3-minimal
Users that are interested in flash-attention-v2-RDNA3-minimal are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AITemplate is a Python framework which renders neural network into high performance CUDA/HIP C++ code. Specialized for FP16 TensorCore (N…☆12Jun 24, 2024Updated 2 years ago
- Fast and memory-efficient exact attention ported to rocm☆14Dec 1, 2023Updated 2 years ago
- The small, fast game engine for Compose Multiplatform☆10Feb 1, 2025Updated last year
- A tiny implementation of in-place FFT. The performance is comparable to FFTW3 for length 2^17 to 2^20.☆15Jul 24, 2018Updated 8 years ago
- Running PyTorch on Windows with AMD GPUs using alpha ROCm wheels. It's fast, it's fragile, and it hates you back.☆23Oct 31, 2025Updated 8 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Image processing tool for ComfyUI☆14Aug 6, 2025Updated 11 months ago
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆34Updated this week
- 8-bit CUDA functions for PyTorch Rocm compatible☆42Mar 26, 2024Updated 2 years ago
- Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.☆15Aug 31, 2023Updated 2 years ago
- Official repository Flash Local Linear Attention☆38May 28, 2026Updated last month
- NES emulator written in pure FreeBASIC with love by Blyss Sarania and Gavin Schulte(Nobbs66).☆21Oct 29, 2025Updated 8 months ago
- GPT2 in handwritten PTX☆15Jun 29, 2025Updated last year
- Decoding Attention is specially optimized for MHA, MQA, GQA and MLA using CUDA core for the decoding stage of LLM inference.☆47Jun 11, 2025Updated last year
- ☆165Sep 15, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Implement FlashAttention v2 with minimal code to learn.☆19Jun 12, 2024Updated 2 years ago
- [DEPRECATED] Moved to ROCm/rocm-libraries repo☆114Updated this week
- ☆72Updated this week
- Standalone Flash Attention v2 kernel without libtorch dependency☆113Sep 10, 2024Updated last year
- RWKV, in easy to read code☆73Mar 25, 2025Updated last year
- Optimized FP16/BF16 x FP4 GPU kernels for AMD GPUs☆63May 29, 2026Updated last month
- Guides to hopefully simplify the process of using ROCm.☆12Sep 26, 2024Updated last year
- ComfyUI custom nodes for DeepSeek, Qwen, GPT, and other OpenAI-compatible LLM APIs, with tools for chat, translation, vision, and JSON wo…☆28Apr 23, 2026Updated 3 months ago
- A convenient fast Text to Speech Whisper Speech by Collabora you can train a voice on the fly on ComfyUI☆45Mar 9, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Quick and easy Diffusers CLI☆15Jun 28, 2026Updated 3 weeks ago
- AutoHotKey script to translate Joystick movement to keypresses.☆12Jun 9, 2014Updated 12 years ago
- GPU monitor for Linux terminal supporting single or multiple gpu's in realtime☆19Updated this week
- Development repository for the Triton language and compiler☆146Updated this week
- A low-cost, high-performance deep learning training framework that enables efficient 100B-scale model fine-tuning on a commodity server w…☆23Mar 21, 2025Updated last year
- YOLOX with NCNN/MNN/TNN/ONNXRuntime C++.☆13Dec 18, 2021Updated 4 years ago
- (ECCV 2026): Official code for Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models☆18Jul 9, 2026Updated 2 weeks ago
- ☆91Jan 23, 2025Updated last year
- A curated list of resources, libraries, tools, and communities for working with Local Large Language Models (LLMs).☆11Dec 20, 2024Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- CUDA on AMD GPUs☆610Feb 11, 2026Updated 5 months ago
- ☆12Feb 7, 2018Updated 8 years ago
- Automated Design of Agentic Systems☆10Sep 7, 2024Updated last year
- ☆49Mar 3, 2024Updated 2 years ago
- A forked version of flux-fast that makes flux-fast even faster with cache-dit, 3.3x speedup on NVIDIA L20.☆24Jul 18, 2025Updated last year
- AI Tensor Engine for ROCm☆502Updated this week
- ☆11Nov 2, 2017Updated 8 years ago