Ongoing research training transformer models at scale
☆43Sep 16, 2026Updated this week
Alternatives and similar repositories for Megatron-LM
Users that are interested in Megatron-LM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MAD (Model Automation and Dashboarding)☆43Updated this week
- ☆76Updated this week
- A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs☆127Updated this week
- A practical guide to high-performance gluon kernel development on AMD GFX9 GPUs.☆51Updated this week
- AI Tensor Engine for ROCm☆565Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆30Sep 10, 2026Updated last week
- Fast and memory-efficient exact attention☆239Aug 12, 2026Updated last month
- A high-performance acceleration library dedicated to large-scale model training on AMD GPUs☆70Updated this week
- Primus-SaFE(Stability and Fault Endurance)☆58Updated this week
- [ICML 2024] SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models☆22May 28, 2024Updated 2 years ago
- Repository with examples and exercises for OLCF and AMD's HIP training series☆17Oct 16, 2023Updated 2 years ago
- ☆12Mar 22, 2022Updated 4 years ago
- a small C++ lattice library☆15Jan 9, 2020Updated 6 years ago
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆94Jul 14, 2026Updated 2 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆62Sep 15, 2023Updated 3 years ago
- Automating analysis from trace files☆91Updated this week
- Scale-out system monitoring☆27Updated this week
- AI Workload Orchestrator for Kubernetes☆25Updated this week
- DEPRECATED REPOSITORY. ROCm Inference Transfer Library (RIXL) is a port of the NIXL library for AMD GPUs. See README_rocm.md for AMD spe…☆15Jun 10, 2026Updated 3 months ago
- A high-throughput and memory-efficient inference and serving engine for LLMs☆124Updated this week
- Documentation for vLLM Dev Channel releases☆10Dec 5, 2024Updated last year
- [DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror☆549Updated this week
- ☆27Oct 9, 2025Updated 11 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [DEPRECATED] Moved to ROCm/rocm-libraries repo☆114Updated this week
- Modular RDMA Interface☆180Updated this week
- Examples illustrating usage of the rocBLAS library☆17Aug 12, 2024Updated 2 years ago
- Hackable and optimized Transformers building blocks, supporting a composable construction.☆34May 29, 2026Updated 3 months ago
- ☆20Sep 8, 2025Updated last year
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆420Updated this week
- CSS-LM: Contrastive Semi-supervised Fine-tuning of Pre-trained Language Models☆11Jul 1, 2023Updated 3 years ago
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆30Updated this week
- R package for unleashing the power of NVIDIA GPU's☆16Jun 4, 2016Updated 10 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud.☆1,597Dec 15, 2025Updated 9 months ago
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆145Sep 10, 2026Updated last week
- python package of rocm-smi-lib☆25Dec 15, 2025Updated 9 months ago
- A ROCm library for GPU-Initiated IO. This provides support for initiating IO from a ROCm-capable GPU against a range of targets including…☆58Updated this week
- Muon in Int8 Precision Made Possible☆20Jun 18, 2026Updated 3 months ago
- Provides the examples to write and build Habana custom kernels using the HabanaTools☆26Apr 15, 2025Updated last year
- Intel Gaudi's Megatron DeepSpeed Large Language Models for training☆18Dec 19, 2024Updated last year