DGXC Benchmarking provides recipes in ready-to-use templates for evaluating performance of specific AI use cases across hardware and software combinations.
☆98Jul 6, 2026Updated 2 weeks ago
Alternatives and similar repositories for dgxc-benchmarking
Users that are interested in dgxc-benchmarking are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- nvloom is a set of tools designed to scalably test MNNVL fabrics.☆50Apr 1, 2026Updated 3 months ago
- NVSentinel is a cross-platform fault remediation service designed to rapidly remediate runtime node-level issues in GPU-accelerated compu…☆346Updated this week
- A TUI-based utility for real-time monitoring of InfiniBand traffic and performance metrics on the local node☆71May 16, 2026Updated 2 months ago
- Dynamo Workshop☆19Nov 7, 2025Updated 8 months ago
- ☆13May 30, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A toolkit for discovering cluster network topology.☆145Updated this week
- ☆51Updated this week
- ☆68Updated this week
- The CSCS ReFrame test suite☆15Updated this week
- ☆26Oct 9, 2025Updated 9 months ago
- A comprehensive Helm chart for monitoring GPU resources in Kubernetes clusters. This tool provides real-time visibility into GPU allocati…☆29Jun 30, 2026Updated 3 weeks ago
- A benchmark framework for Pytorch☆35Mar 14, 2025Updated last year
- Guides and examples to help achieve optimal performance on a NVIDIA Grace CPU☆17Aug 9, 2024Updated last year
- These are lab guides for our customer/partner training and lab environment☆18Apr 15, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- NVIDIA NCCL Tests for Distributed Training☆149Updated this week
- Pavilion is a Python 3 (3.6+) based framework for running and analyzing tests targeting HPC systems.☆46Updated this week
- Benchmarking guide for the Azure AI Infrastructure.☆41Jun 9, 2026Updated last month
- Prometheus collector and exporter for Slurm cluster metrics. A Slinky project.☆18Nov 7, 2025Updated 8 months ago
- A Kubernetes Operator to manage Node OS customizations.☆58Updated this week
- NVIDIA Fleet Intelligence Agent - Host agent for GPU telemetry collection and attestation☆44Updated this week
- Run Slurm on Kubernetes. A Slinky project.☆333Updated this week
- Linux Sysinfo Snapshot☆66Jun 7, 2026Updated last month
- A Streamlit app for exploring available AWS EC2 Capacity Blocks and SageMaker Training Plans across regions and instance types.☆22May 19, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- HPC tests using MPI codes & synthetic benchmarks with IB/RoCE comparisions - from StackHPC Ltd.☆22Jul 11, 2022Updated 4 years ago
- A Slurm-based HPC workload management environment, driven by Ansible.☆72Updated this week
- Tooling for optimized, validated, and reproducible GPU-accelerated AI runtime in Kubernetes☆351Updated this week
- A distributed storage benchmark for file systems, object stores & block devices with support for GPUs☆281Updated this week
- Slurm debian packages☆16Updated this week
- NVIDIA fork of QEMU☆16Jul 15, 2026Updated last week
- Scripts to customize AWS ParallelCluster☆30Jun 11, 2026Updated last month
- A tool for bandwidth measurements on NVIDIA GPUs.☆737Updated this week
- CloudAI Benchmark Framework☆96Updated this week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Python wrappers for the FirecREST API☆12Jun 21, 2026Updated last month
- Prototype of OpenSHMEM for NVIDIA GPUs, developed as part of DoE Design Forward☆24Apr 26, 2018Updated 8 years ago
- NVIDIA Infra Controller - Hardware Lifecycle Management and multitenant networking☆235Updated this week
- Parallel Computing -- Validation Suite: Validation engine for Exascale project benchmarks☆16Mar 26, 2026Updated 3 months ago
- Recipes for reproducing training and serving benchmarks for large machine learning models using GPUs on Google Cloud.☆138Updated this week
- DRA Driver for NVIDIA GPUs☆675Updated this week
- ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage☆94May 12, 2026Updated 2 months ago