NVIDIA Data Center GPU Manager (DCGM) is a project for gathering telemetry and measuring the health of NVIDIA GPUs
☆799Aug 19, 2026Updated last month
Alternatives and similar repositories for DCGM
Users that are interested in DCGM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- NVIDIA GPU metrics exporter for Prometheus leveraging DCGM☆1,893Sep 18, 2026Updated 3 weeks ago
- A tool for bandwidth measurements on NVIDIA GPUs.☆785Jul 28, 2026Updated 2 months ago
- Golang bindings for Nvidia Datacenter GPU Manager (DCGM)☆158Oct 2, 2026Updated last week
- NVIDIA GPU Operator creates, configures, and manages GPUs in Kubernetes☆2,921Updated this week
- NCCL Tests☆1,673Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- NVIDIA Resiliency Extension is a python package for framework developers and users to implement fault-tolerant features. It improves the …☆338Updated this week
- Go Bindings for the NVIDIA Management Library (NVML)☆458Updated this week
- MIG Partition Editor for NVIDIA GPUs☆265Updated this week
- NVIDIA Inference Xfer Library (NIXL)☆1,293Updated this week
- RDMA and SHARP plugins for nccl library☆233Aug 12, 2026Updated last month
- NVSentinel detects and remediates GPU faults on Kubernetes nodes☆394Updated this week
- NVIDIA device plugin for Kubernetes☆3,894Updated this week
- Tools for monitoring NVIDIA GPUs on Linux☆1,074Nov 2, 2021Updated 4 years ago
- Optimized primitives for collective multi-GPU communication☆5,147Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- NVIDIA NCCL Tests for Distributed Training☆160Updated this week
- A fast GPU memory copy library based on NVIDIA GPUDirect RDMA technology☆1,421Sep 15, 2026Updated 3 weeks ago
- CUDA checkpoint and restore utility☆498Jul 6, 2026Updated 3 months ago
- The NVIDIA Driver Manager is a Kubernetes component which assist in seamless upgrades of NVIDIA Driver on each node of the cluster.☆55Oct 2, 2026Updated last week
- ☆402Apr 23, 2024Updated 2 years ago
- NVIDIA GPUDirect Storage Driver☆387Sep 23, 2026Updated 2 weeks ago
- A transparent, in-container GPU resource controller that enforces memory and compute limits by intercepting CUDA calls without applicatio…☆342Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆8,251Updated this week
- GPU plugin to the node feature discovery for Kubernetes☆308May 27, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- DRA Driver for NVIDIA GPUs☆721Updated this week
- Build and run containers leveraging NVIDIA GPUs☆4,592Updated this week
- NVIDIA Network Operator☆373Updated this week
- Tools for building GPU clusters☆1,475Updated this week
- NCCL Profiling Kit☆156Jul 1, 2024Updated 2 years ago
- A CPU+GPU Profiling library that provides access to timeline traces and hardware performance counters.☆997Sep 28, 2026Updated last week
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆6,746Updated this week
- Heterogeneous GPU Sharing on Kubernetes☆4,745Updated this week
- KAI Scheduler is an open source Kubernetes Native scheduler for AI workloads at large scale☆1,559Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆383Updated this week
- The NVIDIA® Tools Extension SDK (NVTX) is a C-based Application Programming Interface (API) for annotating events, code ranges, and resou…☆562Updated this week
- Container plugin for Slurm Workload Manager☆465May 12, 2026Updated 4 months ago
- NVIDIA container runtime library☆1,134Sep 28, 2026Updated last week
- Microsoft Collective Communication Library☆400Aug 25, 2026Updated last month
- MSCCL++: A GPU-driven communication stack for scalable AI applications☆562Updated this week
- The Triton Inference Server provides an optimized cloud and edge inferencing solution.☆11,059Updated this week