A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs
☆119Aug 14, 2026Updated this week
Alternatives and similar repositories for Primus
Users that are interested in Primus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A high-performance acceleration library dedicated to large-scale model training on AMD GPUs☆68Updated this week
- Primus-SaFE(Stability and Fault Endurance)☆58Updated this week
- Automating analysis from trace files☆90Updated this week
- ☆74Updated this week
- Ongoing research training transformer models at scale☆43Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A practical guide to high-performance gluon kernel development on AMD GFX9 GPUs.☆44Jul 31, 2026Updated 2 weeks ago
- AI Tensor Engine for ROCm☆528Updated this week
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆195Aug 8, 2026Updated last week
- DEPRECATED REPOSITORY. ROCm Inference Transfer Library (RIXL) is a port of the NIXL library for AMD GPUs. See README_rocm.md for AMD spe…☆15Jun 10, 2026Updated 2 months ago
- Modular RDMA Interface☆166Updated this week
- amdgpu example code in hip/asm☆67Updated this week
- A PyTorch native platform for training generative AI models☆17Jun 30, 2026Updated last month
- MAD (Model Automation and Dashboarding)☆40Updated this week
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆145Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- AiTer Optimized Model☆156Updated this week
- ☆15Jun 30, 2026Updated last month
- ☆30Aug 5, 2026Updated last week
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆261Updated this week
- Generating Efficient AI-Centric Kernels☆154Updated this week
- MSCCL++: A GPU-driven communication stack for scalable AI applications☆549Updated this week
- Simple experiments on Tenstorrent GraySkull e75 chip☆14Aug 28, 2024Updated last year
- A ROCm library for GPU-Initiated IO. This provides support for initiating IO from a ROCm-capable GPU against a range of targets including…☆56Aug 8, 2026Updated last week
- [EuroSys'25] Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization☆24Apr 13, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror☆542Updated this week
- Unit Scaling demo and experimentation code☆16Mar 12, 2024Updated 2 years ago
- super repo for rocm libraries☆402Updated this week
- Tutorials for NVIDIA CUPTI samples☆72Jul 22, 2026Updated 3 weeks ago
- UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g…☆1,491Updated this week
- Translate Virtual Address To Physical Address in Linux Kernel☆17Dec 27, 2019Updated 6 years ago
- ☆81Updated this week
- ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage☆99May 12, 2026Updated 3 months ago
- A tool for generating information about the matrix multiplication instructions in AMD Radeon™ and AMD Instinct™ accelerators☆143Apr 10, 2026Updated 4 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Nitro-T is a family of text-to-image diffusion models focused on highly efficient training.☆41Jun 4, 2026Updated 2 months ago
- Ahead of Time (AOT) Triton Math Library☆100Updated this week
- Scripts to build AMD ROCm from source.☆16Oct 31, 2024Updated last year
- ☆22Nov 6, 2025Updated 9 months ago
- torchcomms: a modern PyTorch communications API☆388Updated this week
- Agents, and RL environment, for optimizing GPU kernels on AMD ROCm using LLM agents. Benchmarks LLM serving workloads end-to-end, profile…☆76Updated this week
- AMD SMI☆133May 28, 2026Updated 2 months ago