DGXC Benchmarking provides recipes in ready-to-use templates for evaluating performance of specific AI use cases across hardware and software combinations.
☆105Aug 7, 2026Updated this week
Alternatives and similar repositories for dgxc-benchmarking
Users that are interested in dgxc-benchmarking are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- nvloom is a set of tools designed to scalably test MNNVL fabrics.☆57Jul 24, 2026Updated 2 weeks ago
- NVSentinel is a cross-platform fault remediation service designed to rapidly remediate runtime node-level issues in GPU-accelerated compu…☆369Updated this week
- A TUI-based utility for real-time monitoring of InfiniBand traffic and performance metrics on the local node☆72May 16, 2026Updated 2 months ago
- Dynamo Workshop☆19Nov 7, 2025Updated 9 months ago
- ☆13May 30, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A toolkit for discovering cluster network topology.☆154Updated this week
- ☆67Aug 6, 2026Updated last week
- The CSCS ReFrame test suite☆16Updated this week
- ☆26Oct 9, 2025Updated 10 months ago
- A comprehensive Helm chart for monitoring GPU resources in Kubernetes clusters. This tool provides real-time visibility into GPU allocati…☆33Jun 30, 2026Updated last month
- A benchmark framework for Pytorch☆35Mar 14, 2025Updated last year
- Guides and examples to help achieve optimal performance on a NVIDIA Grace CPU☆18Aug 9, 2024Updated 2 years ago
- These are lab guides for our customer/partner training and lab environment☆18Apr 15, 2026Updated 3 months ago
- NVIDIA NCCL Tests for Distributed Training☆154Jul 29, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Pavilion is a Python 3 (3.6+) based framework for running and analyzing tests targeting HPC systems.☆46Updated this week
- Benchmarking guide for the Azure AI Infrastructure.☆41Jun 9, 2026Updated 2 months ago
- Prometheus collector and exporter for Slurm cluster metrics. A Slinky project.☆18Nov 7, 2025Updated 9 months ago
- A Kubernetes Operator to manage Node OS customizations.☆59Updated this week
- Run Slurm on Kubernetes. A Slinky project.☆342Updated this week
- Linux Sysinfo Snapshot☆66Jun 7, 2026Updated 2 months ago
- A Streamlit app for exploring available AWS EC2 Capacity Blocks and SageMaker Training Plans across regions and instance types.☆25Aug 3, 2026Updated last week
- HPC tests using MPI codes & synthetic benchmarks with IB/RoCE comparisions - from StackHPC Ltd.☆22Jul 11, 2022Updated 4 years ago
- A Slurm-based HPC workload management environment, driven by Ansible.☆72Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Tooling for optimized, validated, and reproducible GPU-accelerated AI runtime in Kubernetes☆382Updated this week
- A distributed storage benchmark for file systems, object stores & block devices with support for GPUs☆285Jul 24, 2026Updated 2 weeks ago
- Slurm debian packages☆16Aug 4, 2026Updated last week
- A tool for bandwidth measurements on NVIDIA GPUs.☆748Jul 28, 2026Updated 2 weeks ago
- CloudAI Benchmark Framework☆97Updated this week
- Python wrappers for the FirecREST API☆12Jul 27, 2026Updated 2 weeks ago
- Prototype of OpenSHMEM for NVIDIA GPUs, developed as part of DoE Design Forward☆24Apr 26, 2018Updated 8 years ago
- InfiniBand fabric monitoring daemon written in Go☆32May 22, 2025Updated last year
- Parallel Computing -- Validation Suite: Validation engine for Exascale project benchmarks☆16Mar 26, 2026Updated 4 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Recipes for reproducing training and serving benchmarks for large machine learning models using GPUs on Google Cloud.☆140Updated this week
- DRA Driver for NVIDIA GPUs☆689Updated this week
- ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage☆99May 12, 2026Updated 3 months ago
- A community driven catalog of tools and products that are useful in the world of high performance computing (HPC)☆11Jul 3, 2025Updated last year
- The tool facilitates debugging convergence issues and testing new algorithms and recipes for training LLMs using Nvidia libraries such as…☆21Sep 17, 2025Updated 10 months ago
- ☆37Oct 31, 2025Updated 9 months ago
- Run Slurm as a Kubernetes scheduler. A Slinky project.☆92Aug 6, 2026Updated last week