NVIDIA GPU metrics exporter for Prometheus leveraging DCGM
☆1,856Sep 2, 2026Updated this week
Alternatives and similar repositories for dcgm-exporter
Users that are interested in dcgm-exporter are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- NVIDIA Data Center GPU Manager (DCGM) is a project for gathering telemetry and measuring the health of NVIDIA GPUs☆784Aug 19, 2026Updated 2 weeks ago
- Nvidia GPU exporter for prometheus using nvidia-smi binary OR using NVML☆1,548Updated this week
- NVIDIA GPU Operator creates, configures, and manages GPUs in Kubernetes☆2,857Updated this week
- NVIDIA device plugin for Kubernetes☆3,864Updated this week
- Golang bindings for Nvidia Datacenter GPU Manager (DCGM)☆157Updated this week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Heterogeneous GPU Sharing on Kubernetes☆4,479Updated this week
- DRA Driver for NVIDIA GPUs☆702Updated this week
- Go Bindings for the NVIDIA Management Library (NVML)☆456Aug 23, 2026Updated last week
- KAI Scheduler is an open source Kubernetes Native scheduler for AI workloads at large scale☆1,488Updated this week
- GPU plugin to the node feature discovery for Kubernetes☆308May 27, 2024Updated 2 years ago
- Tools for monitoring NVIDIA GPUs on Linux☆1,074Nov 2, 2021Updated 4 years ago
- Build and run containers leveraging NVIDIA GPUs☆4,536Updated this week
- Exporter for machine metrics☆13,751Updated this week
- Kubernetes-native Job Queueing☆2,932Updated this week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- LeaderWorkerSet: An API for deploying a group of pods as a unit of replication☆806Updated this week
- A Cloud Native Batch System (Project under CNCF)☆5,917Updated this week
- ☆378Updated this week
- Device-plugin for volcano vgpu which support hard resource isolation☆169Updated this week
- A transparent, in-container GPU resource controller that enforces memory and compute limits by intercepting CUDA calls without applicatio…☆330Updated this week
- A toolkit to run Ray applications on Kubernetes☆2,659Updated this week
- Node feature discovery for Kubernetes☆1,072Aug 19, 2026Updated 2 weeks ago
- MIG Partition Editor for NVIDIA GPUs☆262Updated this week
- This is a place for various problem detectors running on the Kubernetes nodes.☆3,456Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A tool for bandwidth measurements on NVIDIA GPUs.☆762Jul 28, 2026Updated last month
- Add-on agent to generate and expose cluster-level metrics.☆6,193Updated this week
- The NVIDIA Driver Manager is a Kubernetes component which assist in seamless upgrades of NVIDIA Driver on each node of the cluster.☆56Updated this week
- Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes☆5,851Updated this week
- NCCL Tests☆1,643Updated this week
- GPU Sharing Scheduler for Kubernetes Cluster☆1,532Dec 29, 2023Updated 2 years ago
- NVIDIA k8s device plugin for Kubevirt☆292Updated this week
- Distributed AI Model Training and LLM Fine-Tuning on Kubernetes☆2,203Updated this week
- ☆906Apr 2, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- NVIDIA Network Operator☆368Updated this week
- Open, Multi-Cloud, Multi-Cluster Kubernetes Orchestration☆5,579Updated this week
- Kubebuilder - SDK for building Kubernetes APIs using CRDs☆9,303Updated this week
- Scalable and efficient source of container resource metrics for Kubernetes built-in autoscaling pipelines.☆6,710Updated this week
- Kubernetes Virtualization API and runtime in order to define and manage virtual machines.☆7,043Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆7,950Updated this week
- NVSentinel detects and remediates GPU faults on Kubernetes nodes☆380Updated this week