GPU environment and cluster management with LLM support
☆660May 16, 2024Updated 2 years ago
Alternatives and similar repositories for genv
Users that are interested in genv are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A top-like tool for monitoring GPUs in a cluster☆85Feb 14, 2024Updated 2 years ago
- Translation layer that maps any Kubernetes framework's Custom Resource Definitions (CRDs) into a standardized, generic structure.☆55Updated this week
- ☆330Updated this week
- KAI Scheduler is an open source Kubernetes Native scheduler for AI workloads at large scale☆1,406Updated this week
- ☆295Jul 5, 2026Updated 2 weeks ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Tensors, for human consumption☆1,390Apr 9, 2026Updated 3 months ago
- ClearML Fractional GPU - Run multiple containers on the same GPU with driver level memory limitation ✨ and compute time-slicing☆93Mar 12, 2026Updated 4 months ago
- ☆811Jul 8, 2026Updated last week
- Run, manage, and scale AI workloads on any AI infrastructure. Use one system to access & manage all AI compute (Kubernetes, Slurm, 20+ cl…☆10,322Updated this week
- Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernete…☆2,187Updated this week
- DRA Driver for NVIDIA GPUs☆675Updated this week
- GPUd automates monitoring, diagnostics, and issue identification for GPUs☆486Updated this week
- Practical GPU Sharing Without Memory Size Constraints☆314Mar 28, 2025Updated last year
- A comprehensive Helm chart for monitoring GPU resources in Kubernetes clusters. This tool provides real-time visibility into GPU allocati…☆29Jun 30, 2026Updated 3 weeks ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- NVIDIA GPU Operator creates, configures, and manages GPUs in Kubernetes☆2,797Updated this week
- Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling☆242Updated this week
- HAMi-core compiles libvgpu.so, which ensures hard limit on GPU in container☆317Jul 10, 2026Updated last week
- A Datacenter Scale Distributed Inference Serving Framework☆7,540Updated this week
- ☆16Updated this week
- GPU plugin to the node feature discovery for Kubernetes☆309May 27, 2024Updated 2 years ago
- Heterogeneous GPU Sharing on Kubernetes☆4,006Updated this week
- LeaderWorkerSet: An API for deploying a group of pods as a unit of replication☆767Updated this week
- A JupyterLab extension for displaying dashboards of GPU usage.☆680Jun 25, 2026Updated 3 weeks ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- The NVIDIA Driver Manager is a Kubernetes component which assist in seamless upgrades of NVIDIA Driver on each node of the cluster.☆55Updated this week
- An Envoy inspired, ultimate LLM-first gateway for LLM serving and downstream application developers and enterprises☆27Apr 24, 2025Updated last year
- The Triton Inference Server provides an optimized cloud and edge inferencing solution.☆10,858Updated this week
- ☆12Aug 27, 2024Updated last year
- ☆207Jul 12, 2026Updated last week
- Supercharge Your Model Training☆5,487Apr 29, 2026Updated 2 months ago
- NVIDIA device plugin for Kubernetes☆3,823Updated this week
- Tools for building GPU clusters☆1,463Updated this week
- Kubernetes Operator for MPI-based applications (distributed training, HPC, etc.)☆530Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- elastic-gpu-scheduler is a Kubernetes scheduler extender for GPU resources scheduling.☆147Nov 21, 2022Updated 3 years ago
- Python client for the Run:ai REST API☆25Dec 15, 2025Updated 7 months ago
- Simple, safe way to store and distribute tensors☆3,811Updated this week
- AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-te…☆1,226Updated this week
- GPU Sharing Scheduler for Kubernetes Cluster☆1,535Dec 29, 2023Updated 2 years ago
- Containers for machine learning☆9,446Updated this week
- A Data Streaming Library for Efficient Neural Network Training☆1,534Jun 25, 2026Updated 3 weeks ago