Optimized FP16/BF16 x FP4 GPU kernels for AMD GPUs
☆69May 29, 2026Updated 3 months ago
Alternatives and similar repositories for petit-kernel
Users that are interested in petit-kernel are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆198Updated this week
- ☆17May 22, 2023Updated 3 years ago
- [NeurIPS'25] FreqExit: Enabling Early-Exit Inference for Visual Autoregressive Models via Frequency-Aware Guidance☆21Dec 15, 2025Updated 8 months ago
- ☆28Jul 2, 2026Updated 2 months ago
- Ongoing research training transformer models at scale☆19Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official implementation of EMNLP'23 paper "Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?"☆24Oct 25, 2023Updated 2 years ago
- Jupyter Notebooks for the Connect Intensive MLND Program☆11Dec 17, 2016Updated 9 years ago
- Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.☆19Feb 9, 2026Updated 6 months ago
- QuickReduce is a performant all-reduce library designed for AMD ROCm that supports inline compression.☆38Aug 29, 2025Updated last year
- AI Tensor Engine for ROCm☆549Updated this week
- The goal of the OSSCI Fleet is to provide a central mechanism to enable test automation, batch job scheduling, and developer access to a …☆13Apr 28, 2026Updated 4 months ago
- a simple Flash Attention v2 implementation with ROCM (RDNA3 GPU, roc wmma), mainly used for stable diffusion(ComfyUI) in Windows ZLUDA en…☆54Aug 25, 2024Updated 2 years ago
- ☆13Jun 16, 2026Updated 2 months ago
- Github mirror of trition-lang/triton repo.☆195Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Automated bottleneck detection and solution orchestration☆22Feb 24, 2026Updated 6 months ago
- FORMULA 2.0: Formal Specifications for Verification and Synthesis☆17May 29, 2024Updated 2 years ago
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆269Updated this week
- ☆22May 30, 2026Updated 3 months ago
- ☆19Mar 29, 2026Updated 5 months ago
- Implementation of NM sparsity recipe presented in the paper "Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers".☆11Feb 5, 2024Updated 2 years ago
- Docker image for the BIRD Internet Routing Daemon☆11May 1, 2021Updated 5 years ago
- Thunder Research Group's Collective Communication Library☆52Jul 8, 2025Updated last year
- NVFP4 Flash-Attention 4 on BlackWell☆53Jul 23, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Fast fp16-fp8 mixed precision matmul on RDNA3/3.5 GPUs without native fp8☆37Updated this week
- AiTer Optimized Model☆168Updated this week
- NetBricks: A network function framework written in Rust and using DPDK☆13Nov 8, 2019Updated 6 years ago
- High-performance LLM operator library built on TileLang.☆176Updated this week
- ☆24May 26, 2026Updated 3 months ago
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆145Aug 10, 2026Updated 3 weeks ago
- My submission for the GPUMODE/AMD fp8 mm challenge☆29Jun 4, 2025Updated last year
- Anatomy of High-Performance GEMM with Online Fault Tolerance on GPUs☆14Apr 3, 2025Updated last year
- Official Code of The Combinatorial Brain Surgeon: Pruning Weights That Cancel One Another in Neural Networks[ICML2022]☆16Sep 20, 2022Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆19Nov 9, 2024Updated last year
- Kata Containers CI☆13Dec 1, 2023Updated 2 years ago
- FlashSampling: Fast and Memory-Efficient Exact Sampling (https://huggingface.co/papers/2603.15854)☆87Aug 5, 2026Updated 3 weeks ago
- ☆155Aug 18, 2025Updated last year
- A lightweight triton-based General Matrix Multiplication (GEMM) library.☆67Jul 21, 2026Updated last month
- ☆48Nov 1, 2025Updated 10 months ago
- ☆76Updated this week