A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs
☆124Sep 6, 2026Updated this week
Alternatives and similar repositories for Primus
Users that are interested in Primus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A high-performance acceleration library dedicated to large-scale model training on AMD GPUs☆69Updated this week
- Primus-SaFE(Stability and Fault Endurance)☆58Updated this week
- ☆75Updated this week
- Ongoing research training transformer models at scale☆43Updated this week
- A practical guide to high-performance gluon kernel development on AMD GFX9 GPUs.☆47Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- AI Tensor Engine for ROCm☆555Updated this week
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆198Updated this week
- DEPRECATED REPOSITORY. ROCm Inference Transfer Library (RIXL) is a port of the NIXL library for AMD GPUs. See README_rocm.md for AMD spe…☆15Jun 10, 2026Updated 2 months ago
- amdgpu example code in hip/asm☆69Aug 10, 2026Updated 3 weeks ago
- Fast and Furious AMD Kernels☆462Aug 27, 2026Updated last week
- Modular RDMA Interface☆174Updated this week
- A PyTorch native platform for training generative AI models☆17Jun 30, 2026Updated 2 months ago
- MAD (Model Automation and Dashboarding)☆43Updated this week
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆145Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- AiTer Optimized Model☆170Updated this week
- Scale-out system monitoring☆27Updated this week
- ☆15Jun 30, 2026Updated 2 months ago
- ☆30Updated this week
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆274Updated this week
- Distributed Compiler and Optimized Parallel Kernels☆1,539Aug 12, 2026Updated 3 weeks ago
- Generating Efficient AI-Centric Kernels☆171Updated this week
- MSCCL++: A GPU-driven communication stack for scalable AI applications☆554Updated this week
- Simple experiments on Tenstorrent GraySkull e75 chip☆14Aug 28, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A ROCm library for GPU-Initiated IO. This provides support for initiating IO from a ROCm-capable GPU against a range of targets including…☆57Aug 8, 2026Updated 3 weeks ago
- [EuroSys'25] Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization☆24Apr 13, 2026Updated 4 months ago
- [DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror☆546Updated this week
- Unit Scaling demo and experimentation code☆15Mar 12, 2024Updated 2 years ago
- super repo for rocm libraries☆419Updated this week
- Tutorials for NVIDIA CUPTI samples☆72Jul 22, 2026Updated last month
- UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g…☆1,507Updated this week
- Translate Virtual Address To Physical Address in Linux Kernel☆17Dec 27, 2019Updated 6 years ago
- ☆83Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage☆102May 12, 2026Updated 3 months ago
- A tool for generating information about the matrix multiplication instructions in AMD Radeon™ and AMD Instinct™ accelerators☆143Apr 10, 2026Updated 4 months ago
- ☆20Updated this week
- Nitro-T is a family of text-to-image diffusion models focused on highly efficient training.☆41Jun 4, 2026Updated 3 months ago
- Ahead of Time (AOT) Triton Math Library☆100Updated this week
- Scripts to build AMD ROCm from source.☆16Oct 31, 2024Updated last year
- ☆22Nov 6, 2025Updated 10 months ago