A high-performance acceleration library dedicated to large-scale model training on AMD GPUs
☆70Sep 23, 2026Updated this week
Alternatives and similar repositories for Primus-Turbo
Users that are interested in Primus-Turbo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Primus-SaFE(Stability and Fault Endurance)☆58Updated this week
- A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs☆128Updated this week
- ☆76Updated this week
- A PyTorch native platform for training generative AI models☆17Jun 30, 2026Updated 2 months ago
- ☆15Jun 30, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Automating analysis from trace files☆91Updated this week
- Ongoing research training transformer models at scale☆43Sep 16, 2026Updated last week
- AI Tensor Engine for ROCm☆567Updated this week
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆202Sep 17, 2026Updated last week
- amdgpu example code in hip/asm☆69Aug 10, 2026Updated last month
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆145Sep 10, 2026Updated 2 weeks ago
- Modular RDMA Interface☆181Updated this week
- MAD (Model Automation and Dashboarding)☆43Updated this week
- [WIP] Better (FP8) attention for Hopper☆33Aug 21, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Training hybrid models for dummies.☆30Nov 1, 2025Updated 10 months ago
- A comprehensive toolkit for streamlining data editing, search, and inspection for large-scale language model training and interpretabilit…☆21Oct 30, 2025Updated 10 months ago
- Fused KL divergence from hidden states for knowledge distillation☆22Apr 28, 2026Updated 4 months ago
- QuickReduce is a performant all-reduce library designed for AMD ROCm that supports inline compression.☆38Aug 29, 2025Updated last year
- mKernel: fast multi-node, multi-GPU fused kernels☆281Updated this week
- An experimental communicating attention kernel based on DeepEP.☆34Jul 29, 2025Updated last year
- NCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.☆25Apr 10, 2026Updated 5 months ago
- A practical guide to high-performance gluon kernel development on AMD GFX9 GPUs.☆51Sep 14, 2026Updated last week
- Generating Efficient AI-Centric Kernels☆181Updated this week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆284Updated this week
- Agents, and RL environment, for optimizing GPU kernels on AMD ROCm using LLM agents. Benchmarks LLM serving workloads end-to-end, profile…☆76Sep 2, 2026Updated 3 weeks ago
- AiTer Optimized Model☆184Updated this week
- Scale-out system monitoring☆27Updated this week
- Notes and artifacts from the ONNX steering committee☆29Sep 13, 2026Updated last week
- ☆24Sep 17, 2026Updated last week
- torchcomms: a modern PyTorch communications API☆397Updated this week
- [DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror☆548Updated this week
- Distributed Compiler and Optimized Parallel Kernels☆1,552Updated this week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- A tool for generating information about the matrix multiplication instructions in AMD Radeon™ and AMD Instinct™ accelerators☆145Apr 10, 2026Updated 5 months ago
- DEPRECATED REPOSITORY. ROCm Inference Transfer Library (RIXL) is a port of the NIXL library for AMD GPUs. See README_rocm.md for AMD spe…☆15Jun 10, 2026Updated 3 months ago
- paNote: an graph note software can be deployed as blog or use as electron☆13Jun 15, 2024Updated 2 years ago
- An experimental CPU backend for Triton (https//github.com/openai/triton)☆48Aug 18, 2025Updated last year
- Fast and memory-efficient exact attention☆239Aug 12, 2026Updated last month
- ☆20Apr 16, 2025Updated last year
- Project showing how to develop NKI kernels for Llama 3.2 1B inference☆21May 29, 2025Updated last year