Power measurement for CUDA programs by polling using NVIDIA Management Library (nvml) APIs.
☆26Jun 24, 2017Updated 9 years ago
Alternatives and similar repositories for nvml-power
Users that are interested in nvml-power are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An llvm pass for counting global uncoalesced acceses for cuda code via dynamic analysis.☆14Nov 17, 2018Updated 7 years ago
- PSTensor provides a way to hack the memory management of tensors in TensorFlow and PyTorch by defining your own C++ Tensor Class.☆10Feb 10, 2022Updated 4 years ago
- A cli to add deploy keys to a repo☆10Jun 29, 2017Updated 9 years ago
- [CF ’20] Verified Instruction-Level Energy Consumption Measurement for NVIDIA GPUs☆15Dec 11, 2020Updated 5 years ago
- A demo project demonstrating the performance improvement by cpp extension, which wrapped with pybind11.☆10Nov 16, 2021Updated 4 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Parse objdump files using tree-sitter☆13Nov 22, 2023Updated 2 years ago
- A parser for PTX 6.5☆13Jun 19, 2023Updated 3 years ago
- CK workflow, portable packages and other artifacts for the ReQuEST-ASPLOS'18 submission:☆11Jan 16, 2019Updated 7 years ago
- ☆11Jun 9, 2023Updated 3 years ago
- ☆12Aug 15, 2023Updated 2 years ago
- SParse AcceleRation on Tensor Architecture☆18Apr 15, 2026Updated 3 months ago
- Time based theme switching for the alacritty terminal☆12Feb 9, 2022Updated 4 years ago
- Logger for MPI communication☆28Jul 12, 2023Updated 3 years ago
- A simple script to plot the Roofline model for given HW platforms and applications☆10Mar 17, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- CK workflow, portable packages and other artifacts for the ReQuEST-ASPLOS'18 submission:☆13Jan 16, 2019Updated 7 years ago
- Example of multi-process, multi-GPU training using Torch-parallel, nVidia-nccl, and nVidia-MPS☆17Sep 22, 2016Updated 9 years ago
- Artifact for OSDI'23: MGG: Accelerating Graph Neural Networks with Fine-grained intra-kernel Communication-Computation Pipelining on Mult…☆40Mar 17, 2024Updated 2 years ago
- Validated Collective Knowledge workflows and results from the 1st ACM ReQuEST tournament on co-design of Pareto-efficient SW/HW stack for…☆13Oct 16, 2018Updated 7 years ago
- Mandelbrot fractal on NVidia GPUs using CUDA dynamic parallelism and Mariani-Silver algorithm☆30Apr 7, 2014Updated 12 years ago
- ibmgraphblas☆29Oct 15, 2018Updated 7 years ago
- A set of cog recipes for C++ reflection☆15Aug 21, 2011Updated 14 years ago
- A parallel programming model for online applications with complex synchronization requirements.☆16Jun 8, 2022Updated 4 years ago
- CK workflow, portable packages and other artifacts for the ReQuEST-ASPLOS'18 submission:☆15Oct 8, 2019Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Emacs major mode for Alloy☆13Jul 14, 2018Updated 8 years ago
- UPP is a minimalist and generic text preprocessor using Lua macros.☆13Oct 13, 2024Updated last year
- Fast Synchronization-Free Algorithms for Parallel Sparse Triangular Solves with Multiple Right-Hand Sides (SpTRSM)☆17Feb 14, 2020Updated 6 years ago
- inference on tvm runtime using c++ with gpu enabled☆10Apr 25, 2018Updated 8 years ago
- Artifact for 'Register Optimizations for Stencils on GPUs'☆10Sep 18, 2018Updated 7 years ago
- Energy Consumption-Aware Tabular Benchmark For Neural Architecture Search☆11Aug 18, 2025Updated 11 months ago
- Seamless integration between Darkman and Emacs using the D-Bus protocol.☆14Nov 12, 2024Updated last year
- ☆11May 25, 2026Updated 2 months ago
- Canopy is a machine learning learning compiler stack with the capability of adopting high-end FPGAs. As a part of OpenAIOS project, Canop…☆12May 7, 2021Updated 5 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Python GUI for differential forms☆13Oct 14, 2023Updated 2 years ago
- Cayley Dickson algebra implementation in python☆13Jan 3, 2019Updated 7 years ago
- Systolic Three Matrix Multiplier for Graph Convolutional Networks using High Level Synthesis☆24Jul 29, 2022Updated 4 years ago
- Crellvm: Verified Credible Compilation for LLVM☆18Jun 26, 2018Updated 8 years ago
- ☆10Dec 8, 2021Updated 4 years ago
- A compact hash algorithm for CPUs and GPUs using OpenCL☆15Sep 26, 2020Updated 5 years ago
- ☆17Nov 13, 2019Updated 6 years ago