A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs
☆108Jul 24, 2026Updated this week
Alternatives and similar repositories for Primus
Users that are interested in Primus are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A high-performance acceleration library dedicated to large-scale model training on AMD GPUs☆67Updated this week
- Primus-SaFE(Stability and Fault Endurance)☆58Updated this week
- Toolkit for launching and observing MaxText training on Slurm-managed GPU clusters☆29Jul 19, 2026Updated last week
- Automating analysis from trace files☆84Updated this week
- ☆72Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Ongoing research training transformer models at scale☆43Updated this week
- A practical guide to high-performance gluon kernel development on AMD GFX9 GPUs.☆41Updated this week
- AI Tensor Engine for ROCm☆503Updated this week
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆193Updated this week
- DEPRECATED REPOSITORY. ROCm Inference Transfer Library (RIXL) is a port of the NIXL library for AMD GPUs. See README_rocm.md for AMD spe…☆15Jun 10, 2026Updated last month
- Modular RDMA Interface☆157Updated this week
- Fast and Furious AMD Kernels☆446Jul 10, 2026Updated 2 weeks ago
- amdgpu example code in hip/asm☆66Updated this week
- A PyTorch native platform for training generative AI models☆17Jun 30, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- MAD (Model Automation and Dashboarding)☆39Updated this week
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆146Updated this week
- AiTer Optimized Model☆144Updated this week
- Scale-out system monitoring☆25Updated this week
- ☆15Jun 30, 2026Updated 3 weeks ago
- ☆30Updated this week
- FlyDSL is the Python front‑end of the project: Flexible LaYout DSL.☆249Updated this week
- Distributed Compiler based on Triton for Parallel Systems☆1,498Updated this week
- Generating Efficient AI-Centric Kernels☆131Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- MSCCL++: A GPU-driven communication stack for scalable AI applications☆542Updated this week
- A ROCm library for GPU-Initiated IO. This provides support for initiating IO from a ROCm-capable GPU against a range of targets including…☆53Jul 10, 2026Updated 2 weeks ago
- [EuroSys'25] Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization☆24Apr 13, 2026Updated 3 months ago
- [DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror☆539Updated this week
- Unit Scaling demo and experimentation code☆16Mar 12, 2024Updated 2 years ago
- super repo for rocm libraries☆390Updated this week
- Tutorials for NVIDIA CUPTI samples☆70Updated this week
- UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g…☆1,471Updated this week
- Translate Virtual Address To Physical Address in Linux Kernel☆17Dec 27, 2019Updated 6 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆81Updated this week
- ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage☆95May 12, 2026Updated 2 months ago
- A tool for generating information about the matrix multiplication instructions in AMD Radeon™ and AMD Instinct™ accelerators☆140Apr 10, 2026Updated 3 months ago
- ☆19Updated this week
- Ahead of Time (AOT) Triton Math Library☆100Updated this week
- Scripts to build AMD ROCm from source.☆16Oct 31, 2024Updated last year
- ☆21Nov 6, 2025Updated 8 months ago