NVSentinel is a cross-platform fault remediation service designed to rapidly remediate runtime node-level issues in GPU-accelerated computing environments
☆373Aug 18, 2026Updated this week
Alternatives and similar repositories for NVSentinel
Users that are interested in NVSentinel are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Tooling for optimized, validated, and reproducible GPU-accelerated AI runtime in Kubernetes☆385Updated this week
- A toolkit for discovering cluster network topology.☆156Updated this week
- A Kubernetes Operator to manage Node OS customizations.☆59Updated this week
- KAI Scheduler is an open source Kubernetes Native scheduler for AI workloads at large scale☆1,453Updated this week
- Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling☆251Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- DGXC Benchmarking provides recipes in ready-to-use templates for evaluating performance of specific AI use cases across hardware and soft…☆106Aug 7, 2026Updated last week
- DRA Driver for NVIDIA GPUs☆691Updated this week
- GPU Cluster Monitoring (GCM): Large-Scale AI Research Cluster Monitoring☆235Updated this week
- Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and i…☆125Updated this week
- GPUd automates monitoring, diagnostics, and issue identification for GPUs☆489Updated this week
- Run Slurm in Kubernetes☆416Updated this week
- LeaderWorkerSet: An API for deploying a group of pods as a unit of replication☆789Updated this week
- NVIDIA Data Center GPU Manager (DCGM) is a project for gathering telemetry and measuring the health of NVIDIA GPUs☆782Jul 6, 2026Updated last month
- NVIDIA Fleet Intelligence Agent - Host agent for GPU telemetry collection and attestation☆49Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- NVIDIA Infra Controller - Hardware Lifecycle Management and multitenant networking☆252Updated this week
- The main purpose of runtime copilot is to assist with node runtime management tasks such as configuring registries, upgrading versions, i…☆13May 16, 2023Updated 3 years ago
- Run Slurm as a Kubernetes scheduler. A Slinky project.☆92Updated this week
- Run Slurm on Kubernetes. A Slinky project.☆346Updated this week
- Platform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.☆198Updated this week
- ☆26Jun 29, 2026Updated last month
- ☆306Aug 3, 2026Updated 2 weeks ago
- Gateway API Inference Extension☆741Updated this week
- NVIDIA NCCL Tests for Distributed Training☆154Updated this week
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- NVIDIA Resiliency Extension is a python package for framework developers and users to implement fault-tolerant features. It improves the …☆322Updated this week
- An Operator for deployment and maintenance of NVIDIA NIMs and NeMo microservices in a Kubernetes environment.☆160Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆7,795Updated this week
- NVIDIA GPU Operator creates, configures, and manages GPUs in Kubernetes☆2,841Updated this week
- nvloom is a set of tools designed to scalably test MNNVL fabrics.☆57Jul 24, 2026Updated 3 weeks ago
- NVIDIA Inference Xfer Library (NIXL)☆1,195Updated this week
- Cloud Native Benchmarking of Foundation Models☆46Jul 31, 2025Updated last year
- Achieve state of the art inference performance with modern accelerators on Kubernetes☆4,049Updated this week
- NVIDIA GPU metrics exporter for Prometheus leveraging DCGM☆1,835Jul 25, 2026Updated 3 weeks ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A comprehensive Helm chart for monitoring GPU resources in Kubernetes clusters. This tool provides real-time visibility into GPU allocati…☆33Jun 30, 2026Updated last month
- An Envoy inspired, ultimate LLM-first gateway for LLM serving and downstream application developers and enterprises☆27Apr 24, 2025Updated last year
- knavigator is a development, testing, and optimization toolkit for AI/ML scheduling systems at scale on Kubernetes.☆80Jul 6, 2026Updated last month
- 🧯 Kubernetes coverage for fault awareness and recovery, works for any LLMOps, MLOps, AI workloads.☆35Updated this week
- DRANET is a Kubernetes Network Driver that uses Dynamic Resource Allocation (DRA) to deliver high-performance networking for demanding ap…☆146Updated this week
- Kubernetes-native Job Queueing☆2,870Updated this week
- Simplified Data Management and Sharing for Kubernetes☆18Updated this week