A high-performance agent for collecting NVIDIA GPU metrics and exporting them via OpenTelemetry protocol.
☆28Jun 24, 2026Updated 3 weeks ago
Alternatives and similar repositories for gpu-metrics-agent
Users that are interested in gpu-metrics-agent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A WIP drop in replacement for Prometheus Alertmanager, built on top of Cloudflare Workers☆25Nov 19, 2025Updated 8 months ago
- This is a tool for managing GPU partitions for NVIDIA Fabric Manager’s Shared NVSwitch.☆18Jul 2, 2026Updated 2 weeks ago
- Prometheus Optimizer enhances Prometheus monitoring by optimizing configuration and rules management.☆36Apr 1, 2025Updated last year
- GitHub Action for Continuous Profiling which you can run to profile your CI/CD. It uses parca and Polar Signals cloud.☆15Feb 10, 2026Updated 5 months ago
- A random PromQL query generator☆25Feb 5, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Leader election for prometheus☆31Dec 18, 2024Updated last year
- Signal Studio is for SREs who refuse noisy observability.☆26Jul 4, 2026Updated 2 weeks ago
- Get insights from your scrape endpoint (even it speaks Protobuf)☆32Mar 19, 2026Updated 4 months ago
- operator that tracks SelinuxPolicy objects in certain namespaces.☆12Jun 26, 2022Updated 4 years ago
- ☆14Jul 16, 2024Updated 2 years ago
- A library to build PromQL expression in Golang.☆32Updated this week
- An eBPF-based tool for detecting and cleaning up stale resources☆30May 13, 2026Updated 2 months ago
- prom-analytics-proxy is a lightweight proxy that captures detailed analytics on PromQL queries, providing valuable insights into query us…☆150Updated this week
- OTLP wire format utilities for Go - zero-allocation counting, sharding, and routing for high-throughput telemetry pipelines☆19Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Create a graph visualization of your Flux Kustomization tree☆19Mar 30, 2026Updated 3 months ago
- High-performance edge proxy for applying telemetry policies before ingestion.☆16Updated this week
- A tool for tracking the usage of metrics across dashboards, alerting rules & recording rules☆52Updated this week
- Kubernetes CNI benchmark 2024-01☆21Apr 16, 2024Updated 2 years ago
- Go libraries for interacting with Hashicorp Vault☆14Jul 13, 2026Updated last week
- Anatomy of High-Performance GEMM with Online Fault Tolerance on GPUs☆14Apr 3, 2025Updated last year
- A layer 2 switch for VMs powered by eBPF☆45Feb 25, 2025Updated last year
- API for coordinating Maintenance in Kubernetes.☆26Jul 2, 2026Updated 2 weeks ago
- Shared library to work with Prometheus and Parquet☆50Jul 13, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Adapter to query Thanos StoreAPI with Prometheus remote read support.☆40Jun 23, 2026Updated 3 weeks ago
- Real-time node group observability for AWS EKS☆24Apr 21, 2025Updated last year
- The official implementation of ICLR26 poster, Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMs.☆18Mar 31, 2026Updated 3 months ago
- Automated Kubernetes workload resource tracking and adjustment☆60Updated this week
- A crate that supports graceful shutdown☆15Jan 22, 2026Updated 5 months ago
- A less opinionated Alertmanager☆14Oct 16, 2023Updated 2 years ago
- Register Cluster-API clusters with Argo-CD☆32Updated this week
- Kubernetes controller to automatically configure Thanos receive hashrings☆112Apr 8, 2026Updated 3 months ago
- A CPU and Hardware Acceleration (GPU) tester for Jellyfin☆10Nov 22, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Tunnel Kubectl requests over `kubectl proxy` to avoid round trips to API server☆11Sep 25, 2024Updated last year
- Performance benchmark of Kubernetes log collectors☆27Mar 22, 2026Updated 3 months ago
- MCP server for LLMs to interact with Prometheus☆65Updated this week
- [EXPERIMENTAL] Manage, troubleshoot and validate Prometheus-Operator resources via Command Line Interface!☆29Updated this week
- Integration of opentelemetry with the tracing crate☆27Jul 1, 2026Updated 2 weeks ago
- A verification tool for ensuring parallelization equivalence in distributed model training.☆17Sep 1, 2025Updated 10 months ago
- ITIX's Custom CoreOS build☆10Nov 24, 2020Updated 5 years ago