Dynolog is a telemetry daemon for performance monitoring and tracing. It exports metrics from different components in the system like the linux kernel, CPU, disks, Intel PT, GPUs etc. Dynolog also integrates with pytorch and can trigger traces for distributed training applications.
☆380Aug 7, 2026Updated this week
Alternatives and similar repositories for dynolog
Users that are interested in dynolog are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A library to analyze PyTorch traces.☆543May 29, 2026Updated 2 months ago
- A CPU+GPU Profiling library that provides access to timeline traces and hardware performance counters.☆987Updated this week
- PArametrized Recommendation and Ai Model benchmark is a repository for development of numerous uBenchmarks as well as end to end nets for…☆155Jul 2, 2026Updated last month
- Meta's fleetwide profiler framework☆350Jul 7, 2026Updated last month
- NCCL Profiling Kit☆156Jul 1, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- GPUd automates monitoring, diagnostics, and issue identification for GPUs☆487Updated this week
- CUDA checkpoint and restore utility☆480Jul 6, 2026Updated last month
- ☆26Jun 29, 2026Updated last month
- Collection of scripts to build PyTorch and the domain libraries from source.☆14Jul 9, 2026Updated last month
- BDC is the eBPF powered DNS caching mechanism in kernel inspired by BMC☆10May 13, 2022Updated 4 years ago
- Fault tolerance for PyTorch (HSDP, LocalSGD, DiLoCo, Streaming DiLoCo)☆530Jul 16, 2026Updated 3 weeks ago
- [OSDI'24] Serving LLM-based Applications Efficiently with Semantic Variable☆223Sep 21, 2024Updated last year
- NVIDIA Inference Xfer Library (NIXL)☆1,184Updated this week
- MSLK (Meta Superintelligence Labs Kernels) is a collection of PyTorch GPU operator libraries that are designed and optimized for GenAI tr…☆145Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- CUPTI based GPU profiling library exposing usdt hooks☆37Jun 30, 2026Updated last month
- Lightweight daemon for monitoring CUDA runtime API calls with eBPF uprobes☆155Mar 29, 2025Updated last year
- ☆18May 16, 2022Updated 4 years ago
- ☆21Jul 7, 2026Updated last month
- GPU-CR: GPU Checkpoint & Restore☆26Jun 4, 2026Updated 2 months ago
- A tool for bandwidth measurements on NVIDIA GPUs.☆745Jul 28, 2026Updated last week
- CUDA Kernel Benchmarking Library☆915Updated this week
- A Datacenter Scale Distributed Inference Serving Framework☆7,713Updated this week
- NVIDIA Data Center GPU Manager (DCGM) is a project for gathering telemetry and measuring the health of NVIDIA GPUs☆776Jul 6, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Microsoft Collective Communication Library☆394Sep 20, 2023Updated 2 years ago
- TransferBench is a utility capable of benchmarking simultaneous copies between user-specified devices (CPUs/GPUs)☆76Updated this week
- MSCCL++: A GPU-driven communication stack for scalable AI applications☆546Updated this week
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆153May 28, 2026Updated 2 months ago
- Collective communications library with various primitives for multi-machine training.☆1,440Jul 31, 2026Updated last week
- Scripts to identify an application that will benefit from code locality optimization on ARM architecture and to generate an optimized lin…☆19Jan 26, 2024Updated 2 years ago
- The NVIDIA® Tools Extension SDK (NVTX) is a C-based Application Programming Interface (API) for annotating events, code ranges, and resou…☆551Updated this week
- ☆13Feb 6, 2026Updated 6 months ago
- A low-latency & high-throughput serving engine for LLMs☆516Jan 8, 2026Updated 7 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- TritonParse: A Compiler Tracer, Visualizer, and Reproducer for Triton Kernels☆213Updated this week
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆6,208Updated this week
- Fast OS-level support for GPU checkpoint and restore☆288Sep 28, 2025Updated 10 months ago
- Userspace eBPF runtime for Observability, Network, GPU & General Extensions Framework☆1,546Updated this week
- FlashInfer: Kernel Library for LLM Serving☆6,133Updated this week
- Byted PyTorch Distributed for Hyperscale Training of LLMs and RLs☆1,036Mar 3, 2026Updated 5 months ago
- LLTFI is a tool, which is an extension of LLFI, allowing users to run fault injection experiments on C/C++, TensorFlow and PyTorch applic…☆45Jul 9, 2026Updated last month