nvloom is a set of tools designed to scalably test MNNVL fabrics.
☆65Jul 24, 2026Updated last month
Alternatives and similar repositories for nvloom
Users that are interested in nvloom are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The CSCS ReFrame test suite☆17Updated this week
- Exemplar Performance provides recipes in ready-to-use templates for evaluating performance of specific AI use cases across hardware and s…☆112Aug 26, 2026Updated 3 weeks ago
- ☆23Sep 6, 2026Updated 2 weeks ago
- Kerberos ticket delegation and impersonation for Batch/CI/CD environments☆22Jul 28, 2026Updated last month
- A service-aware RoCE network monitoring system based on end- to-end probing.☆34Aug 24, 2026Updated 3 weeks ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆38Jan 19, 2023Updated 3 years ago
- Using C++ templates to track dimensional metadata☆11Nov 20, 2020Updated 5 years ago
- PathwaysJob API is an OSS Kubernetes-native API, to deploy ML training and batch inference workloads, using Pathways on GKE.☆23Jul 9, 2026Updated 2 months ago
- Python wrappers for the FirecREST API☆12Updated this week
- InfiniBand fabric monitoring daemon written in Go☆32May 22, 2025Updated last year
- ☆38Oct 31, 2025Updated 10 months ago
- A tool for bandwidth measurements on NVIDIA GPUs.☆777Jul 28, 2026Updated last month
- Tools to deploy GPU clusters in the Cloud☆34Apr 4, 2023Updated 3 years ago
- Benchmarking guide for the Azure AI Infrastructure.☆39Jun 9, 2026Updated 3 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A NCCL extension library, designed to efficiently offload GPU memory allocated by the NCCL communication library.☆119Dec 17, 2025Updated 9 months ago
- General interest repository for CSCS users☆52Feb 20, 2025Updated last year
- Optiflop measures the optimally achievable FLOPs for mathematical operations on various platforms.☆14Jul 29, 2026Updated last month
- A toolkit for discovering cluster network topology.☆170Updated this week
- A Feishu/Lark AI agent bot☆16Feb 27, 2026Updated 6 months ago
- NVIDIA NCCL Tests for Distributed Training☆159Updated this week
- ☆27Jun 29, 2026Updated 2 months ago
- ☆19Jan 17, 2025Updated last year
- Pavilion is a Python 3 (3.6+) based framework for running and analyzing tests targeting HPC systems.☆46Sep 10, 2026Updated last week
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Aims to implement dual-port and multi-qp solutions in deepEP ibrc transport☆75May 9, 2025Updated last year
- CloudAI Benchmark Framework☆99Updated this week
- A simple GPU reservation tool for single host shared development systems☆30Aug 24, 2026Updated 3 weeks ago
- Parallel Computing -- Validation Suite: Validation engine for Exascale project benchmarks☆16Mar 26, 2026Updated 5 months ago
- NVSentinel detects and remediates GPU faults on Kubernetes nodes☆385Updated this week
- ☆11Apr 10, 2019Updated 7 years ago
- Scripts for managing Debian and RPM package repositories☆17Jan 14, 2026Updated 8 months ago
- Communication patterns for AI, built on top of NCCL device and host APIs☆58Updated this week
- ⭐️ Simple package browsing portal. ⭐️☆19Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Proof of concept of view_maybe☆12Dec 9, 2024Updated last year
- Converts an Infiniband topology file to graphviz dot format or slurm topology.conf format☆18Feb 2, 2026Updated 7 months ago
- An MPI ABI compatibility layer☆34Jun 17, 2026Updated 3 months ago
- ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage☆103May 12, 2026Updated 4 months ago
- Terraform examples for deploying HPC clusters on OCI☆70Sep 8, 2026Updated last week
- NVIDIA Performance Libraries: Sample code☆23May 28, 2026Updated 3 months ago
- NVIDIA Fleet Intelligence Agent - Host agent for GPU telemetry collection and attestation☆53Updated this week