A high-performance acceleration library dedicated to large-scale model training on AMD GPUs
☆67Jul 24, 2026Updated this week
Alternatives and similar repositories for Primus-Turbo
Users that are interested in Primus-Turbo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Primus-SaFE(Stability and Fault Endurance)☆58Updated this week
- A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs☆108Updated this week
- Toolkit for launching and observing MaxText training on Slurm-managed GPU clusters☆29Jul 19, 2026Updated last week
- ☆72Updated this week
- A PyTorch native platform for training generative AI models☆17Jun 30, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆15Jun 30, 2026Updated 3 weeks ago
- Automating analysis from trace files☆84Updated this week
- Ongoing research training transformer models at scale☆43Updated this week
- AI Tensor Engine for ROCm☆503Updated this week
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆193Updated this week
- amdgpu example code in hip/asm☆66Updated this week
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆146Updated this week
- Modular RDMA Interface☆157Updated this week
- MAD (Model Automation and Dashboarding)☆39Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [WIP] Better (FP8) attention for Hopper☆33Feb 24, 2025Updated last year
- Training hybrid models for dummies.☆31Nov 1, 2025Updated 8 months ago
- Fast and Furious AMD Kernels☆446Jul 10, 2026Updated 2 weeks ago
- A comprehensive toolkit for streamlining data editing, search, and inspection for large-scale language model training and interpretabilit…☆21Oct 30, 2025Updated 8 months ago
- QuickReduce is a performant all-reduce library designed for AMD ROCm that supports inline compression.☆38Aug 29, 2025Updated 10 months ago
- mKernel: fast multi-node, multi-GPU fused kernels☆255Jun 21, 2026Updated last month
- An experimental communicating attention kernel based on DeepEP.☆34Jul 29, 2025Updated 11 months ago
- NCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.☆23Apr 10, 2026Updated 3 months ago
- FlyDSL is the Python front‑end of the project: Flexible LaYout DSL.☆249Updated this week
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Generating Efficient AI-Centric Kernels☆131Updated this week
- A practical guide to high-performance gluon kernel development on AMD GFX9 GPUs.☆41Updated this week
- Agents, and RL environment, for optimizing GPU kernels on AMD ROCm using LLM agents. Benchmarks LLM serving workloads end-to-end, profile…☆71Updated this week
- AiTer Optimized Model☆144Updated this week
- Notes and artifacts from the ONNX steering committee☆29Updated this week
- Scale-out system monitoring☆25Updated this week
- ☆24May 26, 2026Updated 2 months ago
- torchcomms: a modern PyTorch communications API☆380Updated this week
- [DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror☆539Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- An implementation of model parallel autoregressive transformers on GPUs, based on the DeepSpeed library.☆21Nov 28, 2022Updated 3 years ago
- The goal of the OSSCI Fleet is to provide a central mechanism to enable test automation, batch job scheduling, and developer access to a …☆13Apr 28, 2026Updated 2 months ago
- Distributed Compiler based on Triton for Parallel Systems☆1,498Updated this week
- A tool for generating information about the matrix multiplication instructions in AMD Radeon™ and AMD Instinct™ accelerators☆140Apr 10, 2026Updated 3 months ago
- DEPRECATED REPOSITORY. ROCm Inference Transfer Library (RIXL) is a port of the NIXL library for AMD GPUs. See README_rocm.md for AMD spe…☆15Jun 10, 2026Updated last month
- An experimental CPU backend for Triton (https//github.com/openai/triton)☆48Aug 18, 2025Updated 11 months ago
- Fast and memory-efficient exact attention☆234Jul 16, 2026Updated last week