A comprehensive Helm chart for monitoring GPU resources in Kubernetes clusters. This tool provides real-time visibility into GPU allocation, utilization, memory usage, and pod status through an integrated Prometheus and Grafana monitoring stack.
☆34Aug 26, 2026Updated last month
Alternatives and similar repositories for gpu-usage-monitor
Users that are interested in gpu-usage-monitor are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆21Jul 28, 2026Updated 2 months ago
- ☆326Updated this week
- Simulate NVIDIA infrastructure (e.g. GPU) on CPU nodes in Kubernetes ️☆90Updated this week
- NVIDIA Fleet Command is a hybrid-cloud platform for securely and remotely deploying, managing, and scaling AI across dozens or up to thou…☆17Jul 20, 2022Updated 4 years ago
- Identify and reduce instances of underutilization by the users of high-performance computing systems☆18Sep 23, 2026Updated 2 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- dra-driver for HAMi☆26Sep 29, 2026Updated last week
- CUDA keyring packaging for Debian☆14Apr 14, 2023Updated 3 years ago
- CSI driver for CernVM-FS☆26Oct 2, 2026Updated last week
- A toolkit for discovering cluster network topology.☆184Updated this week
- A collection of useful Go libraries for use with NVIDIA GPU management tools☆58Updated this week
- Workflow based on github issues.☆11Apr 30, 2019Updated 7 years ago
- A federation scheduler for multi-cluster☆78Mar 6, 2026Updated 7 months ago
- Translation layer that maps any Kubernetes framework's Custom Resource Definitions (CRDs) into a standardized, generic structure.☆75Updated this week
- KAI Scheduler is an open source Kubernetes Native scheduler for AI workloads at large scale☆1,559Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A community-maintained website about the CashTokens technology, including technical specifications, documentation, guides, and other reso…☆20Oct 14, 2025Updated 11 months ago
- An Operator for deployment and maintenance of NVIDIA NIMs and NeMo microservices in a Kubernetes environment.☆160Updated this week
- Rate Limit operator for Envoy Proxy☆25Nov 15, 2023Updated 2 years ago
- ☆12Oct 1, 2026Updated last week
- AWS Deep Learning Containers (DLCs) are a set of Docker images for training and serving models in TensorFlow, TensorFlow 2, PyTorch, and …☆13Feb 11, 2025Updated last year
- Ochami deployment recipes☆14Jul 16, 2026Updated 2 months ago
- NVIDIA CPU microcode☆13Mar 10, 2015Updated 11 years ago
- 将原项目改造成go module项目。GoReporter Use Go Modules.☆13Sep 7, 2021Updated 5 years ago
- Holodeck is a project to create test environments optimised for GPU projects.☆30Sep 3, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Platform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.☆230Updated this week
- ☆22Mar 19, 2025Updated last year
- AkashOS is an unattended install of Ubuntu Server that will become the operating system of the machine. Akash OS will create a Kubernetes…☆13Dec 22, 2025Updated 9 months ago
- Tegra scripts☆13Mar 30, 2017Updated 9 years ago
- FFmpeg-enabled Python Video☆13Mar 23, 2022Updated 4 years ago
- Open OnDemand core library☆18Oct 2, 2026Updated last week
- A top-like tool for monitoring GPUs in a cluster☆85Feb 14, 2024Updated 2 years ago
- NVSentinel detects and remediates GPU faults on Kubernetes nodes☆394Updated this week
- ☆13Jan 14, 2026Updated 8 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling☆273Updated this week
- A benchmark to evaluate search-augmented LLMs☆17Aug 28, 2025Updated last year
- Đây là repository chứa các tìm hiểu về OpenStack của nhóm Cloud Bách Khoa Hà Nội.☆10Mar 25, 2022Updated 4 years ago
- Description: source code of the OPTEE_OS for NVIDIA Jetson Linux☆12Jan 12, 2023Updated 3 years ago
- Community Helm Charts provided by Offchain Labs☆12Updated this week
- ☆12Oct 28, 2019Updated 6 years ago
- Official Tinkerbell Roadmap☆11Jun 12, 2025Updated last year