A practical guide to high-performance gluon kernel development on AMD GFX9 GPUs.
☆42Jul 31, 2026Updated this week
Alternatives and similar repositories for gfx950-gluon-tutorials
Users that are interested in gfx950-gluon-tutorials are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆74Updated this week
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆257Updated this week
- A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs☆118Updated this week
- Ongoing research training transformer models at scale☆43Updated this week
- ☆30Jul 24, 2026Updated last week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A tool for generating information about the matrix multiplication instructions in AMD Radeon™ and AMD Instinct™ accelerators☆140Apr 10, 2026Updated 3 months ago
- DEPRECATED REPOSITORY. ROCm Inference Transfer Library (RIXL) is a port of the NIXL library for AMD GPUs. See README_rocm.md for AMD spe…☆15Jun 10, 2026Updated last month
- OpenVINO LLM Benchmark☆11Dec 7, 2023Updated 2 years ago
- ☆62Jul 16, 2026Updated 2 weeks ago
- Modular RDMA Interface☆164Updated this week
- MAD (Model Automation and Dashboarding)☆40Updated this week
- Bandwidth test for ROCm☆86Jul 17, 2026Updated 2 weeks ago
- Fast and Furious AMD Kernels☆449Updated this week
- Ahead of Time (AOT) Triton Math Library☆100Jul 23, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Automating analysis from trace files☆86Updated this week
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆194Updated this week
- SYCL materials for ENCCS workshop☆25Apr 25, 2023Updated 3 years ago
- ☆13Jul 21, 2026Updated 2 weeks ago
- Github mirror of trition-lang/triton repo.☆183Updated this week
- OpenMP offload playground☆10Nov 16, 2024Updated last year
- An implemention of parallel marching cubes algorithm by CUDA☆10Sep 23, 2021Updated 4 years ago
- Artifact for 'Register Optimizations for Stencils on GPUs'☆10Sep 18, 2018Updated 7 years ago
- ☆10May 20, 2022Updated 4 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Finite Field Operations on GPGPU☆15Jul 23, 2023Updated 3 years ago
- Materials for workshop on GPU computation for statistics, data science, machine learning applications.☆14Sep 8, 2016Updated 9 years ago
- Gray-Scott reaction-diffusion system in 3D using CUDA☆12Jun 8, 2019Updated 7 years ago
- ☆11Nov 16, 2024Updated last year
- hipDF - GPU DataFrame Library☆20Updated this week
- Cahn Hilliard CUDA (Phase-Field Simulation of Spinodal Decomposition)☆13Jul 4, 2019Updated 7 years ago
- ASTER 💫 : Assembly Tooling and Representations☆34Jul 1, 2026Updated last month
- [DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror☆540Updated this week
- Chapel-based Optimization☆18Jul 21, 2026Updated last week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Unit Scaling demo and experimentation code☆16Mar 12, 2024Updated 2 years ago
- Primus-SaFE(Stability and Fault Endurance)☆58Updated this week
- AI Tensor Engine for ROCm☆513Updated this week
- AiTer Optimized Model☆148Updated this week
- ☆14Jun 30, 2026Updated last month
- ☆11Jun 29, 2021Updated 5 years ago
- ☆15Jul 15, 2023Updated 3 years ago