CUDA checkpoint and restore utility
☆474Jul 6, 2026Updated 2 weeks ago
Alternatives and similar repositories for cuda-checkpoint
Users that are interested in cuda-checkpoint are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- cricket is a virtualization solution for GPUs☆248Sep 9, 2025Updated 10 months ago
- A tool for coordinated checkpoint/restore of distributed applications with CRIU☆35May 11, 2026Updated 2 months ago
- Fast OS-level support for GPU checkpoint and restore☆286Sep 28, 2025Updated 9 months ago
- Orchestrated process and container checkpointing☆133Updated this week
- Hooked CUDA-related dynamic libraries by using automated code generation tools.☆173Dec 12, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆21Jul 10, 2025Updated last year
- NCCL Profiling Kit☆155Jul 1, 2024Updated 2 years ago
- Checkpoint/Restore tool☆3,920Updated this week
- NVIDIA Resiliency Extension is a python package for framework developers and users to implement fault-tolerant features. It improves the …☆310Updated this week
- HAMi-core compiles libvgpu.so, which ensures hard limit on GPU in container☆317Jul 10, 2026Updated last week
- ☆47Dec 13, 2024Updated last year
- Artifacts for our NSDI'23 paper TGS☆97Jun 10, 2024Updated 2 years ago
- NVIDIA Data Center GPU Manager (DCGM) is a project for gathering telemetry and measuring the health of NVIDIA GPUs☆763Jul 6, 2026Updated 2 weeks ago
- A tool for bandwidth measurements on NVIDIA GPUs.☆734Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Fault tolerance for PyTorch (HSDP, LocalSGD, DiLoCo, Streaming DiLoCo)☆523Updated this week
- Checkpoint and Restore in Kubernetes☆168May 15, 2024Updated 2 years ago
- ☆20Jul 7, 2026Updated 2 weeks ago
- A fast GPU memory copy library based on NVIDIA GPUDirect RDMA technology☆1,399Updated this week
- DRA Driver for NVIDIA GPUs☆674Updated this week
- Go Bindings for the NVIDIA Management Library (NVML)☆447Updated this week
- NCCL Tests☆1,595Jul 9, 2026Updated last week
- GPUd automates monitoring, diagnostics, and issue identification for GPUs☆486Updated this week
- DLRover: An Automatic Distributed Deep Learning System☆1,672Jul 6, 2026Updated 2 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- NVIDIA Inference Xfer Library (NIXL)☆1,139Updated this week
- A toolkit for discovering cluster network topology.☆144Updated this week
- GLake: optimizing GPU memory management and IO transmission.☆501Mar 24, 2025Updated last year
- ☆544Jun 7, 2024Updated 2 years ago
- ☆330Updated this week
- Module, Model, and Tensor Serialization/Deserialization☆318Jul 7, 2026Updated last week
- NVIDIA's launch, startup, and logging scripts used by our MLPerf Training and HPC submissions☆43May 13, 2026Updated 2 months ago
- Go Bindings for CRIU☆245Updated this week
- Scripts for managing a large H100 cluster and fixing hardware issues to ensure smooth model training.☆322Aug 20, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- This is a tool for managing GPU partitions for NVIDIA Fabric Manager’s Shared NVSwitch.☆17Jul 2, 2026Updated 2 weeks ago
- TiledLower is a Dataflow Analysis and Codegen Framework written in Rust.☆13Nov 23, 2024Updated last year
- Optimized primitives for collective multi-GPU communication☆11May 8, 2024Updated 2 years ago
- NVIDIA GPUDirect Storage Driver☆367Jun 1, 2026Updated last month
- A Datacenter Scale Distributed Inference Serving Framework☆7,540Updated this week
- Microsoft Collective Communication Library☆394Sep 20, 2023Updated 2 years ago
- The NVIDIA Driver Manager is a Kubernetes component which assist in seamless upgrades of NVIDIA Driver on each node of the cluster.☆55Updated this week