Fast and Furious AMD Kernels
☆444Jul 10, 2026Updated last week
Alternatives and similar repositories for HipKittens
Users that are interested in HipKittens are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- amdgpu example code in hip/asm☆66Jul 9, 2026Updated last week
- AI Tensor Engine for ROCm☆497Updated this week
- FlyDSL is the Python front‑end of the project: Flexible LaYout DSL.☆237Updated this week
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆193Updated this week
- AiTer Optimized Model☆141Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Modular RDMA Interface☆151Updated this week
- A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs☆107Updated this week
- Tile primitives for speedy kernels☆3,552Jul 13, 2026Updated last week
- Efficient implementation of DeepSeek Ops (Blockwise FP8 GEMM, MoE, and MLA) for AMD Instinct MI300X☆79Feb 11, 2026Updated 5 months ago
- A practical guide to high-performance gluon kernel development on AMD GFX9 GPUs.☆38Updated this week
- Super fast FP32 matrix multiplication on RDNA3☆92Mar 30, 2025Updated last year
- A lightweight triton-based General Matrix Multiplication (GEMM) library.☆65Jun 13, 2026Updated last month
- A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.☆910Updated this week
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)☆1,148Mar 24, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror☆538Updated this week
- A Quirky Assortment of CuTe Kernels☆1,063Updated this week
- Automating analysis from trace files☆82Updated this week
- Distributed Compiler based on Triton for Parallel Systems☆1,494Updated this week
- Kernels, of the mega variety :)☆780May 26, 2026Updated last month
- ☆61Updated this week
- Mirage Persistent Kernel: Compiling LLMs into a MegaKernel☆2,376Updated this week
- ☆15Jun 10, 2022Updated 4 years ago
- ☆24May 26, 2026Updated last month
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆30Jun 16, 2026Updated last month
- CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs☆230Updated this week
- Github mirror of trition-lang/triton repo.☆178Updated this week
- ☆65Apr 26, 2025Updated last year
- A high-performance acceleration library dedicated to large-scale model training on AMD GPUs☆67Updated this week
- ☆72Updated this week
- mKernel: fast multi-node, multi-GPU fused kernels☆251Jun 21, 2026Updated last month
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels☆6,674Updated this week
- Accelerating MoE with IO and Tile-aware Optimizations☆732Jul 4, 2026Updated 2 weeks ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Tilus is a tile-level kernel programming language with explicit control over shared memory and registers.☆489Jul 5, 2026Updated 2 weeks ago
- My submission for the GPUMODE/AMD fp8 mm challenge☆29Jun 4, 2025Updated last year
- TileFusion is an experimental C++ macro kernel template library that elevates the abstraction level in CUDA C for tile processing.☆115Jun 28, 2025Updated last year
- FlashInfer: Kernel Library for LLM Serving☆5,988Updated this week
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆146Jul 15, 2026Updated last week
- This repo contains data and code for the paper "Reasoning over Public and Private Data in Retrieval-Based Systems."☆47Jul 18, 2024Updated 2 years ago
- CUDA Templates and Python DSLs for High-Performance Linear Algebra☆10,104Updated this week