A practical guide to high-performance gluon kernel development on AMD GFX9 GPUs.
☆48Sep 11, 2026Updated this week
Alternatives and similar repositories for gfx950-gluon-tutorials
Users that are interested in gfx950-gluon-tutorials are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆75Updated this week
- ☆18Apr 10, 2026Updated 5 months ago
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆279Updated this week
- Ongoing research training transformer models at scale☆43Updated this week
- ☆30Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- amdgpu example code in hip/asm☆69Aug 10, 2026Updated last month
- A tool for generating information about the matrix multiplication instructions in AMD Radeon™ and AMD Instinct™ accelerators☆143Apr 10, 2026Updated 5 months ago
- DEPRECATED REPOSITORY. ROCm Inference Transfer Library (RIXL) is a port of the NIXL library for AMD GPUs. See README_rocm.md for AMD spe…☆15Jun 10, 2026Updated 3 months ago
- OpenVINO LLM Benchmark☆11Dec 7, 2023Updated 2 years ago
- ☆73Updated this week
- Modular RDMA Interface☆178Updated this week
- MAD (Model Automation and Dashboarding)☆43Updated this week
- Bandwidth test for ROCm☆88Jul 17, 2026Updated last month
- Fast and Furious AMD Kernels☆463Aug 27, 2026Updated 2 weeks ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Ahead of Time (AOT) Triton Math Library☆100Updated this week
- Automating analysis from trace files☆90Updated this week
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆201Updated this week
- SYCL materials for ENCCS workshop☆25Apr 25, 2023Updated 3 years ago
- ☆13Jul 21, 2026Updated last month
- Github mirror of trition-lang/triton repo.☆195Updated this week
- OpenMP offload playground☆10Nov 16, 2024Updated last year
- Computes the Henry coefficient of methane in IRMOF-1☆10Oct 5, 2021Updated 4 years ago
- An implemention of parallel marching cubes algorithm by CUDA☆10Sep 23, 2021Updated 4 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Artifact for 'Register Optimizations for Stencils on GPUs'☆10Sep 18, 2018Updated 7 years ago
- ☆10May 20, 2022Updated 4 years ago
- Finite Field Operations on GPGPU☆15Jul 23, 2023Updated 3 years ago
- Materials for workshop on GPU computation for statistics, data science, machine learning applications.☆14Sep 8, 2016Updated 10 years ago
- Gray-Scott reaction-diffusion system in 3D using CUDA☆12Jun 8, 2019Updated 7 years ago
- ☆11Nov 16, 2024Updated last year
- hipDF - GPU DataFrame Library☆20Updated this week
- Cahn Hilliard CUDA (Phase-Field Simulation of Spinodal Decomposition)☆13Jul 4, 2019Updated 7 years ago
- Semi-automated OpenVINO benchmark_app with variable parameters. User can specify multiple options for any parameters in the benchmark_app…☆10Apr 19, 2022Updated 4 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror☆548Updated this week
- Chapel-based Optimization☆19Jul 21, 2026Updated last month
- ASTER 💫 : Assembly Tooling and Representations☆34Jul 1, 2026Updated 2 months ago
- Primus-SaFE(Stability and Fault Endurance)☆58Updated this week
- AI Tensor Engine for ROCm☆561Updated this week
- AiTer Optimized Model☆175Updated this week
- ☆15Jun 30, 2026Updated 2 months ago