π A curated list of awesome matrix-matrix multiplication (A * B = C) frameworks, libraries and software
β68Feb 23, 2025Updated last year
Alternatives and similar repositories for awesome-gemm
Users that are interested in awesome-gemm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FLA but cuTileβ27Apr 17, 2026Updated 4 months ago
- A curated collection of resources, tutorials, and best practices for learning and mastering NVIDIA CUTLASSβ270May 6, 2025Updated last year
- A Toy-Purpose TPU Simulatorβ24Jun 7, 2024Updated 2 years ago
- High Performance FP8 GEMM Kernels for SM89 and later GPUs.β21Jan 24, 2025Updated last year
- FlashTile is a CUDA Tile IR compiler that is compatible with NVIDIA's tileiras, targeting SM70 through SM121 NVIDIA GPUs.β61Feb 6, 2026Updated 6 months ago
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A Triton JIT runtime and ffi provider in C++β39Aug 7, 2026Updated last week
- CUDA 8-bit Tensor Core Matrix Multiplication based on m16n16k16 WMMA APIβ37Sep 15, 2023Updated 2 years ago
- NCCL Examples from Official NVIDIA NCCL Developer Guide.β21May 29, 2018Updated 8 years ago
- Lab assignments for the Agile Hardware Design courseβ19Nov 14, 2025Updated 9 months ago
- Decoding Attention is specially optimized for MHA, MQA, GQA and MLA using CUDA core for the decoding stage of LLM inference.β48Jun 11, 2025Updated last year
- Fast CUDA matrix multiplication from scratchβ1,285Sep 2, 2025Updated 11 months ago
- β26May 9, 2025Updated last year
- CarND Semantic Segmentationβ16Apr 1, 2018Updated 8 years ago
- GPT2 in handwritten PTXβ16Jun 29, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- β‘οΈQwen-Image 4.8xπ speedup with Hybrid Acceleration for low VRAM GPUsβ17Oct 24, 2025Updated 9 months ago
- Automated bottleneck detection and solution orchestrationβ23Feb 24, 2026Updated 5 months ago
- πππ This repository lists some awesome public CUDA, cuda-python, cuBLAS, cuDNN, CUTLASS, TensorRT, TensorRT-LLM, Triton, TVM, MLIR, PTβ¦β512Aug 2, 2025Updated last year
- Synthesis using Synopsys DC and Physical Design flow using Synopsys ICC II, of my RISC-V 5 stage pipelined using 32 nm tech repoβ15Jul 31, 2024Updated 2 years ago
- A repository where GPU applications are aggregated using a common build flow that supports multiple CUDA versions.β95Apr 14, 2026Updated 4 months ago
- β22May 30, 2026Updated 2 months ago
- β122May 16, 2025Updated last year
- A parser for PTX 6.5β13Jun 19, 2023Updated 3 years ago
- A demonstrative example of running SGLang Diffusion with DP routerβ19Mar 15, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Flash Attention from Scratch on CUDA Ampereβ191Sep 1, 2025Updated 11 months ago
- Acceleration codes for the Ozaki-scheme on integer matrix multiplication units.β27Dec 10, 2025Updated 8 months ago
- Public benchmark results from Kernel Arena, a leaderboard for LLM-generated AI accelerator kernels.β21Mar 11, 2026Updated 5 months ago
- β11Jun 9, 2023Updated 3 years ago
- From Minimal GEMM to Everythingβ234Jul 9, 2026Updated last month
- High-Performance FP32 GEMM on CUDA devicesβ126Jan 21, 2025Updated last year
- Evaluating Large Language Models for CUDA Code Generation ComputeEval is a framework designed to generate and evaluate CUDA code from Larβ¦β144Updated this week
- Anatomy of High-Performance GEMM with Online Fault Tolerance on GPUsβ14Apr 3, 2025Updated last year
- LLM Inference via Triton (Flexible & Modular): Focused on Kernel Optimization using CUBIN binaries, Starting from gpt-oss Modelβ119Apr 28, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- This repository contains the results and code for the MLPerfβ’ Training v1.1 benchmark.β23May 18, 2023Updated 3 years ago
- This serves as a repository for reproducibility of the SC21 paper "In-Depth Analyses of Unified Virtual Memory System for GPU Acceleratedβ¦β37Sep 25, 2023Updated 2 years ago
- FB (Facebook) + GEMM (General Matrix-Matrix Multiplication) - https://code.fb.com/ml-applications/fbgemm/β1,583Updated this week
- Automated CUDA kernel performance diagnostics from NVIDIA Nsight Compute (NCU) CSV exports.β35Mar 18, 2026Updated 5 months ago
- Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.β15Aug 31, 2023Updated 2 years ago
- β21Oct 17, 2025Updated 10 months ago
- Applied AI experiments and examples for PyTorchβ323Aug 22, 2025Updated 11 months ago