GPUd automates monitoring, diagnostics, and issue identification for GPUs
☆489Aug 18, 2026Updated this week
Alternatives and similar repositories for gpud
Users that are interested in gpud are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Scripts for managing a large H100 cluster and fixing hardware issues to ensure smooth model training.☆324Aug 20, 2024Updated last year
- NVSentinel is a cross-platform fault remediation service designed to rapidly remediate runtime node-level issues in GPU-accelerated compu…☆373Updated this week
- KAI Scheduler is an open source Kubernetes Native scheduler for AI workloads at large scale☆1,453Updated this week
- knavigator is a development, testing, and optimization toolkit for AI/ML scheduling systems at scale on Kubernetes.☆80Jul 6, 2026Updated last month
- A tool to detect infrastructure issues on cloud native AI systems☆55Sep 18, 2025Updated 11 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Fault tolerance for PyTorch (HSDP, LocalSGD, DiLoCo, Streaming DiLoCo)☆531Jul 16, 2026Updated last month
- Kubernetes-native Job Queueing☆2,870Updated this week
- InfiniBand fabric monitoring daemon written in Go☆32May 22, 2025Updated last year
- Dynolog is a telemetry daemon for performance monitoring and tracing. It exports metrics from different components in the system like the…☆380Updated this week
- NVIDIA NCCL Tests for Distributed Training☆154Updated this week
- DRA Driver for NVIDIA GPUs☆691Updated this week
- NVIDIA GPU Operator creates, configures, and manages GPUs in Kubernetes☆2,841Updated this week
- NVIDIA Resiliency Extension is a python package for framework developers and users to implement fault-tolerant features. It improves the …☆322Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆7,795Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Heterogeneous GPU Sharing on Kubernetes☆4,383Updated this week
- A toolkit for discovering cluster network topology.☆156Updated this week