A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs
☆131Sep 27, 2026Updated this week
Alternatives and similar repositories for Primus
Users that are interested in Primus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A high-performance acceleration library dedicated to large-scale model training on AMD GPUs☆70Updated this week
- Primus-SaFE(Stability and Fault Endurance)☆58Updated this week
- Toolkit for launching and observing MaxText training on Slurm-managed GPU clusters☆29Jul 19, 2026Updated 2 months ago
- Automating analysis from trace files☆91Updated this week
- ☆76Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Ongoing research training transformer models at scale☆43Updated this week
- A practical guide to high-performance gluon kernel development on AMD GFX9 GPUs.☆51Updated this week
- AI Tensor Engine for ROCm☆571Updated this week
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆203Updated this week
- Fast and Furious AMD Kernels☆471Updated this week
- Modular RDMA Interface☆181Updated this week
- A PyTorch native platform for training generative AI models☆17Jun 30, 2026Updated 2 months ago
- MAD (Model Automation and Dashboarding)☆43Updated this week
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆145Sep 10, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Scale-out system monitoring☆27Updated this week
- AiTer Optimized Model☆187Updated this week
- ☆15Updated this week
- ☆30Updated this week
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆286Updated this week
- Distributed Compiler and Optimized Parallel Kernels☆1,552Sep 18, 2026Updated last week
- Generating Efficient AI-Centric Kernels☆181Updated this week
- MSCCL++: A GPU-driven communication stack for scalable AI applications☆560Updated this week
- Simple experiments on Tenstorrent GraySkull e75 chip☆14Aug 28, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A ROCm library for GPU-Initiated IO. This provides support for initiating IO from a ROCm-capable GPU against a range of targets including…☆60Updated this week
- [EuroSys'25] Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization☆24Apr 13, 2026Updated 5 months ago
- [DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror☆547Updated this week
- Unit Scaling demo and experimentation code☆15Mar 12, 2024Updated 2 years ago
- super repo for rocm libraries☆439Updated this week
- Tutorials for NVIDIA CUPTI samples☆73Jul 22, 2026Updated 2 months ago
- UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g…☆1,528Sep 20, 2026Updated last week
- Translate Virtual Address To Physical Address in Linux Kernel☆17Dec 27, 2019Updated 6 years ago
- ☆84Updated this week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage☆104May 12, 2026Updated 4 months ago
- ☆20Updated this week
- A tool for generating information about the matrix multiplication instructions in AMD Radeon™ and AMD Instinct™ accelerators☆145Updated this week
- Nitro-T is a family of text-to-image diffusion models focused on highly efficient training.☆41Jun 4, 2026Updated 3 months ago
- Ahead of Time (AOT) Triton Math Library☆101Updated this week
- Scripts to build AMD ROCm from source.☆16Oct 31, 2024Updated last year
- torchcomms: a modern PyTorch communications API☆398Updated this week