NVIDIA GPU metrics exporter for Prometheus leveraging DCGM
☆1,811Jul 15, 2026Updated last week
Alternatives and similar repositories for dcgm-exporter
Users that are interested in dcgm-exporter are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- NVIDIA Data Center GPU Manager (DCGM) is a project for gathering telemetry and measuring the health of NVIDIA GPUs☆768Jul 6, 2026Updated 2 weeks ago
- Nvidia GPU exporter for prometheus using nvidia-smi binary☆1,514Updated this week
- NVIDIA GPU Operator creates, configures, and manages GPUs in Kubernetes☆2,806Updated this week
- NVIDIA device plugin for Kubernetes☆3,827Updated this week
- Golang bindings for Nvidia Datacenter GPU Manager (DCGM)☆156Jul 8, 2026Updated 2 weeks ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Heterogeneous GPU Sharing on Kubernetes☆4,058Updated this week
- DRA Driver for NVIDIA GPUs☆677Updated this week
- Go Bindings for the NVIDIA Management Library (NVML)☆447Jul 16, 2026Updated last week
- KAI Scheduler is an open source Kubernetes Native scheduler for AI workloads at large scale☆1,409Updated this week
- GPU plugin to the node feature discovery for Kubernetes☆309May 27, 2024Updated 2 years ago
- Tools for monitoring NVIDIA GPUs on Linux☆1,075Nov 2, 2021Updated 4 years ago
- Build and run containers leveraging NVIDIA GPUs☆4,482Updated this week
- Exporter for machine metrics☆13,638Updated this week
- LeaderWorkerSet: An API for deploying a group of pods as a unit of replication☆769Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Kubernetes-native Job Queueing☆2,745Updated this week
- A Cloud Native Batch System (Project under CNCF)☆5,807Updated this week
- ☆375Updated this week
- Device-plugin for volcano vgpu which support hard resource isolation☆162Jun 9, 2026Updated last month
- HAMi-core compiles libvgpu.so, which ensures hard limit on GPU in container☆319Updated this week
- A toolkit to run Ray applications on Kubernetes☆2,606Updated this week
- Node feature discovery for Kubernetes☆1,056Updated this week
- MIG Partition Editor for NVIDIA GPUs☆259Updated this week
- This is a place for various problem detectors running on the Kubernetes nodes.☆3,438Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A tool for bandwidth measurements on NVIDIA GPUs.☆737Updated this week
- Add-on agent to generate and expose cluster-level metrics.☆6,159Updated this week
- The NVIDIA Driver Manager is a Kubernetes component which assist in seamless upgrades of NVIDIA Driver on each node of the cluster.☆55Updated this week
- Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes☆5,730Updated this week
- NCCL Tests☆1,603Jul 9, 2026Updated 2 weeks ago
- GPU Sharing Scheduler for Kubernetes Cluster☆1,534Dec 29, 2023Updated 2 years ago
- NVIDIA k8s device plugin for Kubevirt☆288Jul 15, 2026Updated last week
- Distributed AI Model Training and LLM Fine-Tuning on Kubernetes☆2,153Updated this week
- ☆904Apr 2, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- NVIDIA Network Operator☆357Updated this week
- Open, Multi-Cloud, Multi-Cluster Kubernetes Orchestration☆5,541Updated this week
- Kubebuilder - SDK for building Kubernetes APIs using CRDs☆9,259Updated this week
- Scalable and efficient source of container resource metrics for Kubernetes built-in autoscaling pipelines.☆6,683Updated this week
- Kubernetes Virtualization API and runtime in order to define and manage virtual machines.☆6,963Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆7,574Updated this week
- NVSentinel is a cross-platform fault remediation service designed to rapidly remediate runtime node-level issues in GPU-accelerated compu…☆350Updated this week