a simple Flash Attention v2 implementation with ROCM (RDNA3 GPU, roc wmma), mainly used for stable diffusion(ComfyUI) in Windows ZLUDA environments.
☆54Aug 25, 2024Updated 2 years ago
Alternatives and similar repositories for flash-attention-v2-RDNA3-minimal
Users that are interested in flash-attention-v2-RDNA3-minimal are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Simple monkeypatch to boost AMD Navi 3 GPUs☆51Apr 21, 2025Updated last year
- AITemplate is a Python framework which renders neural network into high performance CUDA/HIP C++ code. Specialized for FP16 TensorCore (N…☆12Jun 24, 2024Updated 2 years ago
- Fast and memory-efficient exact attention ported to rocm☆14Dec 1, 2023Updated 2 years ago
- The small, fast game engine for Compose Multiplatform☆10Aug 10, 2026Updated last month
- Running PyTorch on Windows with AMD GPUs using alpha ROCm wheels. It's fast, it's fragile, and it hates you back.☆23Oct 31, 2025Updated 10 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Image processing tool for ComfyUI☆14Sep 16, 2026Updated last week
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆37Sep 13, 2026Updated last week
- LoRAFusion: Efficient LoRA Fine-Tuning for LLMs☆30Jul 2, 2026Updated 2 months ago
- 8-bit CUDA functions for PyTorch Rocm compatible☆42Mar 26, 2024Updated 2 years ago
- Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.☆15Aug 31, 2023Updated 3 years ago
- Official repository Flash Local Linear Attention☆40May 28, 2026Updated 3 months ago
- NES emulator written in pure FreeBASIC with love by Blyss Sarania and Gavin Schulte(Nobbs66).☆21Oct 29, 2025Updated 10 months ago
- Flash Attention in raw Cuda C beating PyTorch☆39May 14, 2024Updated 2 years ago
- GPT2 in handwritten PTX☆16Jun 29, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Everything you need to setup on your AMD system for Machine Learning Stuff☆19Jul 31, 2025Updated last year
- Decoding Attention is specially optimized for MHA, MQA, GQA and MLA using CUDA core for the decoding stage of LLM inference.☆48Jun 11, 2025Updated last year
- Fast and memory-efficient exact attention☆239Aug 12, 2026Updated last month
- ☆165Sep 15, 2023Updated 3 years ago
- Implement FlashAttention v2 with minimal code to learn.☆19Jun 12, 2024Updated 2 years ago
- [DEPRECATED] Moved to ROCm/rocm-libraries repo☆114Updated this week
- ☆76Updated this week
- Standalone Flash Attention v2 kernel without libtorch dependency☆113Sep 10, 2024Updated 2 years ago
- RWKV, in easy to read code☆74Mar 25, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Optimized FP16/BF16 x FP4 GPU kernels for AMD GPUs☆70Sep 5, 2026Updated 2 weeks ago
- Installation script for an AI applications using ROCm on Linux.☆50Sep 5, 2026Updated 2 weeks ago
- ComfyUI custom nodes for DeepSeek, Qwen, GPT, and other OpenAI-compatible LLM APIs, with tools for chat, translation, vision, and JSON wo…☆28Apr 23, 2026Updated 5 months ago
- Quick and easy Diffusers CLI☆15Updated this week
- AutoHotKey script to translate Joystick movement to keypresses.☆12Jun 9, 2014Updated 12 years ago
- GPU monitor for Linux terminal supporting single or multiple gpu's in realtime☆20Jul 19, 2026Updated 2 months ago
- A low-cost, high-performance deep learning training framework that enables efficient 100B-scale model fine-tuning on a commodity server w…☆23Mar 21, 2025Updated last year
- YOLOX with NCNN/MNN/TNN/ONNXRuntime C++.☆13Dec 18, 2021Updated 4 years ago
- ComfyUI custom nodes for RVC related inference and image generation☆39Oct 15, 2025Updated 11 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆92Jan 23, 2025Updated last year
- 🤖 Telegram chatbot frontend for Searx.☆16Nov 25, 2018Updated 7 years ago
- (ECCV 2026): Official code for Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models☆23Jul 9, 2026Updated 2 months ago
- GPU implementation of Winograd convolution☆10Oct 23, 2017Updated 8 years ago
- This project is about convolution operator optimization on GPU, include GEMM based (Implicit GEMM) convolution.☆44Sep 29, 2025Updated 11 months ago
- ☆12Feb 7, 2018Updated 8 years ago
- Automated Design of Agentic Systems☆10Sep 7, 2024Updated 2 years ago