Optimized FP16/BF16 x FP4 GPU kernels for AMD GPUs
☆64May 29, 2026Updated 2 months ago
Alternatives and similar repositories for petit-kernel
Users that are interested in petit-kernel are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- End-to-End Super Resolution Object Detection Networks☆12Jun 8, 2018Updated 8 years ago
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆195Updated this week
- ☆17May 22, 2023Updated 3 years ago
- ☆28Jul 2, 2026Updated last month
- Ongoing research training transformer models at scale☆19Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Official implementation of EMNLP'23 paper "Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?"☆24Oct 25, 2023Updated 2 years ago
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 6 months ago
- Repo for Udacity Machine Learning Nanodegree☆10Feb 15, 2017Updated 9 years ago
- The goal of the OSSCI Fleet is to provide a central mechanism to enable test automation, batch job scheduling, and developer access to a …☆13Apr 28, 2026Updated 3 months ago
- a simple Flash Attention v2 implementation with ROCM (RDNA3 GPU, roc wmma), mainly used for stable diffusion(ComfyUI) in Windows ZLUDA en…☆54Aug 25, 2024Updated last year
- Github mirror of trition-lang/triton repo.☆186Updated this week
- Automated bottleneck detection and solution orchestration☆23Feb 24, 2026Updated 5 months ago
- ☆28Jun 18, 2026Updated last month
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆261Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆20May 30, 2026Updated 2 months ago
- ☆19Mar 29, 2026Updated 4 months ago
- Implementation of NM sparsity recipe presented in the paper "Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers".☆11Feb 5, 2024Updated 2 years ago
- Thunder Research Group's Collective Communication Library☆53Jul 8, 2025Updated last year
- NVFP4 Flash-Attention 4 on BlackWell☆39Jul 23, 2026Updated 3 weeks ago
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆34Updated this week
- AiTer Optimized Model☆156Updated this week
- High-performance LLM operator library built on TileLang.☆168Updated this week
- ☆24May 26, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆145Updated this week
- Anatomy of High-Performance GEMM with Online Fault Tolerance on GPUs☆14Apr 3, 2025Updated last year
- Official Code of The Combinatorial Brain Surgeon: Pruning Weights That Cancel One Another in Neural Networks[ICML2022]☆16Sep 20, 2022Updated 3 years ago
- ☆19Nov 9, 2024Updated last year
- FlashSampling: Fast and Memory-Efficient Exact Sampling (https://huggingface.co/papers/2603.15854)☆86Aug 5, 2026Updated last week
- ASTER 💫 : Assembly Tooling and Representations☆34Jul 1, 2026Updated last month
- ☆150Aug 18, 2025Updated 11 months ago
- A lightweight triton-based General Matrix Multiplication (GEMM) library.☆67Jul 21, 2026Updated 3 weeks ago
- ☆48Nov 1, 2025Updated 9 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆74Updated this week
- Sample Codes using NVSHMEM on Multi-GPU☆30Jan 22, 2023Updated 3 years ago
- Modular RDMA Interface☆166Updated this week
- A dynamic binary instrumentation tool for tracing and analyzing CUDA kernel instructions.☆78Updated this week
- [CVPR 2024] DiffAgent: Fast and Accurate Text-to-Image API Selection with Large Language Model☆19Apr 16, 2024Updated 2 years ago
- [ACL 2026 🔥] CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark☆37Apr 20, 2026Updated 3 months ago
- [ICLR 2026] Offical implementation of "OBS-Diff".☆67Mar 5, 2026Updated 5 months ago