NVIDIA GPU metrics exporter for Prometheus leveraging DCGM
☆1,879Sep 18, 2026Updated this week
Alternatives and similar repositories for dcgm-exporter
Users that are interested in dcgm-exporter are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- NVIDIA Data Center GPU Manager (DCGM) is a project for gathering telemetry and measuring the health of NVIDIA GPUs☆794Aug 19, 2026Updated last month
- Nvidia GPU exporter for prometheus using nvidia-smi binary OR using NVML☆1,556Updated this week
- NVIDIA GPU Operator creates, configures, and manages GPUs in Kubernetes☆2,880Updated this week
- NVIDIA device plugin for Kubernetes☆3,877Updated this week
- Golang bindings for Nvidia Datacenter GPU Manager (DCGM)☆158Sep 9, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Heterogeneous GPU Sharing on Kubernetes☆4,659Updated this week
- DRA Driver for NVIDIA GPUs☆711Updated this week
- Go Bindings for the NVIDIA Management Library (NVML)☆457Updated this week
- KAI Scheduler is an open source Kubernetes Native scheduler for AI workloads at large scale☆1,530Updated this week
- GPU plugin to the node feature discovery for Kubernetes☆308May 27, 2024Updated 2 years ago
- Tools for monitoring NVIDIA GPUs on Linux☆1,075Nov 2, 2021Updated 4 years ago
- Build and run containers leveraging NVIDIA GPUs☆4,580Updated this week
- Exporter for machine metrics☆13,797Updated this week
- Kubernetes-native Job Queueing☆2,996Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- LeaderWorkerSet: An API for deploying a group of pods as a unit of replication☆819Updated this week
- A Cloud Native Batch System (Project under CNCF)☆5,976Updated this week
- ☆383Sep 16, 2026Updated last week
- Device-plugin for volcano vgpu which support hard resource isolation☆172Updated this week
- A transparent, in-container GPU resource controller that enforces memory and compute limits by intercepting CUDA calls without applicatio…☆335Updated this week
- A toolkit to run Ray applications on Kubernetes☆2,701Updated this week
- Node feature discovery for Kubernetes☆1,075Updated this week
- MIG Partition Editor for NVIDIA GPUs☆264Updated this week
- This is a place for various problem detectors running on the Kubernetes nodes.☆3,464Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A tool for bandwidth measurements on NVIDIA GPUs.☆779Jul 28, 2026Updated last month
- Add-on agent to generate and expose cluster-level metrics.☆6,206Updated this week
- The NVIDIA Driver Manager is a Kubernetes component which assist in seamless upgrades of NVIDIA Driver on each node of the cluster.☆55Updated this week
- Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes☆5,988Updated this week
- NCCL Tests☆1,663Aug 28, 2026Updated 3 weeks ago
- GPU Sharing Scheduler for Kubernetes Cluster☆1,532Dec 29, 2023Updated 2 years ago
- Distributed AI Model Training and LLM Fine-Tuning on Kubernetes☆2,226Updated this week
- NVIDIA k8s device plugin for Kubevirt☆294Updated this week
- ☆907Apr 2, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- NVIDIA Network Operator☆371Updated this week
- Open, Multi-Cloud, Multi-Cluster Kubernetes Orchestration☆5,692Updated this week
- Kubebuilder - SDK for building Kubernetes APIs using CRDs☆9,319Updated this week
- Scalable and efficient source of container resource metrics for Kubernetes built-in autoscaling pipelines.☆6,731Updated this week
- Kubernetes Virtualization API and runtime in order to define and manage virtual machines.☆7,084Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆8,145Updated this week
- NVSentinel detects and remediates GPU faults on Kubernetes nodes☆387Updated this week