GPU Cluster Monitoring (GCM): Large-Scale AI Research Cluster Monitoring
☆238Aug 31, 2026Updated 2 weeks ago
Alternatives and similar repositories for gcm
Users that are interested in gcm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Clusterscope is a CLI and python library to extract information from HPC Clusters and Jobs.☆24Jul 14, 2026Updated 2 months ago
- NVSentinel detects and remediates GPU faults on Kubernetes nodes☆385Updated this week
- Fault tolerance for PyTorch (HSDP, LocalSGD, DiLoCo, Streaming DiLoCo)☆540Aug 28, 2026Updated 3 weeks ago
- ☆33Apr 19, 2025Updated last year
- PyTorch Single Controller☆1,076Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A Top-Down Profiler for GPU Applications☆24Feb 29, 2024Updated 2 years ago
- NVIDIA NCCL Tests for Distributed Training☆159Updated this week
- Exemplar Performance provides recipes in ready-to-use templates for evaluating performance of specific AI use cases across hardware and s…☆112Aug 26, 2026Updated 3 weeks ago
- Where GPUs get cooked 👩🍳🔥☆407Updated this week
- NVIDIA Resiliency Extension is a python package for framework developers and users to implement fault-tolerant features. It improves the …☆337Updated this week
- Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200…☆1,733Updated this week
- Run Slurm as a Kubernetes scheduler. A Slinky project.☆100Updated this week
- torchcomms: a modern PyTorch communications API☆396Updated this week
- Run Slurm on Kubernetes. A Slinky project.☆363Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- htop-like TUI for real-time RDMA network monitoring.☆104Updated this week
- PCCL (Prime Collective Communications Library) implements fault tolerant collective communications over IP☆159Sep 12, 2025Updated last year
- Accelerating MoE with IO and Tile-aware Optimizations☆769Aug 29, 2026Updated 3 weeks ago
- Real-time terminal monitor for InfiniBand networks - htop for high-speed interconnects☆142Sep 3, 2026Updated 2 weeks ago
- Lightweight harness for replaying inference traffic against an endpoint☆28Aug 1, 2026Updated last month