A high-performance acceleration library dedicated to large-scale model training on AMD GPUs
☆69Sep 2, 2026Updated this week
Alternatives and similar repositories for Primus-Turbo
Users that are interested in Primus-Turbo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Primus-SaFE(Stability and Fault Endurance)☆58Updated this week
- A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs☆123Updated this week
- ☆76Updated this week
- A PyTorch native platform for training generative AI models☆17Jun 30, 2026Updated 2 months ago
- ☆15Jun 30, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Automating analysis from trace files☆90Updated this week
- Ongoing research training transformer models at scale☆43Updated this week
- AI Tensor Engine for ROCm☆551Updated this week
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆198Updated this week
- amdgpu example code in hip/asm☆69Aug 10, 2026Updated 3 weeks ago
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆145Updated this week
- Modular RDMA Interface☆174Updated this week
- MAD (Model Automation and Dashboarding)☆42Updated this week
- [WIP] Better (FP8) attention for Hopper☆33Aug 21, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Training hybrid models for dummies.☆30Nov 1, 2025Updated 10 months ago
- Fast and Furious AMD Kernels☆462Aug 27, 2026Updated last week
- A comprehensive toolkit for streamlining data editing, search, and inspection for large-scale language model training and interpretabilit…☆21Oct 30, 2025Updated 10 months ago
- Fused KL divergence from hidden states for knowledge distillation☆22Apr 28, 2026Updated 4 months ago
- QuickReduce is a performant all-reduce library designed for AMD ROCm that supports inline compression.☆38Aug 29, 2025Updated last year
- mKernel: fast multi-node, multi-GPU fused kernels☆270Updated this week
- An experimental communicating attention kernel based on DeepEP.☆34Jul 29, 2025Updated last year
- NCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.☆25Apr 10, 2026Updated 4 months ago
- Generating Efficient AI-Centric Kernels☆167Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆270Updated this week
- A practical guide to high-performance gluon kernel development on AMD GFX9 GPUs.☆46Updated this week
- Agents, and RL environment, for optimizing GPU kernels on AMD ROCm using LLM agents. Benchmarks LLM serving workloads end-to-end, profile…☆76Updated this week
- AiTer Optimized Model☆169Updated this week
- Notes and artifacts from the ONNX steering committee☆29Aug 26, 2026Updated last week
- torchcomms: a modern PyTorch communications API☆391Updated this week
- Learning a game engine by example.☆10Feb 8, 2016Updated 10 years ago
- ☆12Jan 17, 2025Updated last year
- The simplest, fastest repository for training/finetuning medium-sized GPTs.☆38Dec 3, 2023Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- GPU & cluster health and performance monitoring solution for OCI☆15Mar 12, 2026Updated 5 months ago
- [DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror☆545Updated this week
- An implementation of model parallel autoregressive transformers on GPUs, based on the DeepSpeed library.☆22Nov 28, 2022Updated 3 years ago
- Distributed Compiler and Optimized Parallel Kernels☆1,532Aug 12, 2026Updated 3 weeks ago
- A tool for generating information about the matrix multiplication instructions in AMD Radeon™ and AMD Instinct™ accelerators☆143Apr 10, 2026Updated 4 months ago
- DEPRECATED REPOSITORY. ROCm Inference Transfer Library (RIXL) is a port of the NIXL library for AMD GPUs. See README_rocm.md for AMD spe…☆15Jun 10, 2026Updated 2 months ago
- An experimental CPU backend for Triton (https//github.com/openai/triton)☆48Aug 18, 2025Updated last year