Light-weight Performance Variance Detection for Production-run Parallel Applications
☆16Aug 28, 2023Updated 2 years ago
Alternatives and similar repositories for VAPRO
Users that are interested in VAPRO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A GPU FP32 computation method with Tensor Cores.☆27Dec 8, 2025Updated 8 months ago
- C-Coupler2: a flexible and user-friendly community coupler for model coupling and nesting☆41Sep 4, 2019Updated 6 years ago
- PerFlow-AI is a programmable performance analysis, modeling, prediction tool for AI system.☆33May 12, 2026Updated 3 months ago
- code for examining determinism of performance counters☆22Mar 18, 2021Updated 5 years ago
- A portable and efficient infrastracture for value profilers. Doc: https://vclinic.readthedocs.io/en/latest/index.html☆14Mar 4, 2026Updated 5 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- examples for jupyterlab porjects☆11Dec 8, 2022Updated 3 years ago
- Fortran IO Netcdf Assembly☆19Sep 12, 2021Updated 4 years ago
- Democratizing AlphaFold3: an PyTorch reimplementation to accelerate protein structure prediction☆22May 24, 2025Updated last year
- FlipIt: An LLVM Based Fault Injector for HPC☆15May 14, 2021Updated 5 years ago
- 西电操作系统课设避坑指南☆10Sep 7, 2020Updated 5 years ago
- The code for our paper "Neural Architecture Search as Program Transformation Exploration"☆17Apr 28, 2021Updated 5 years ago
- Searching for a Strategy: Modelling Player Trajectories in Soccer Games using Social LSTM☆16Dec 20, 2017Updated 8 years ago
- A system for programming formally-verified loop transformations.☆16Jan 31, 2019Updated 7 years ago
- Test cases for MIPS CPU implementation☆12Dec 26, 2019Updated 6 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Sequence-level 1F1B schedule for LLMs.☆19Jun 4, 2024Updated 2 years ago
- NVIDIA DPU OPs collection☆15Mar 6, 2023Updated 3 years ago
- Sample code and application to simplifying onboarding new hosts to the network with DNA Center☆14Dec 8, 2022Updated 3 years ago
- ☆13Jan 23, 2021Updated 5 years ago
- NVIDIA GPU direct RDMA using SISCI API☆18Apr 8, 2018Updated 8 years ago
- Official implementation of Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores.☆18Nov 13, 2025Updated 9 months ago
- Paper: "Aggregating Capacity in FL through Successive Layer Training for Computationally-Constrained Devices"☆17Jan 10, 2024Updated 2 years ago
- ☆20Jul 7, 2017Updated 9 years ago
- ☆23Mar 31, 2012Updated 14 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Einsum optimization using opt_einsum and PyTorch FX graph rewriting☆22Mar 17, 2022Updated 4 years ago
- Tools and library to manipulate EFI variables.☆10Apr 21, 2026Updated 3 months ago
- Extending the HDF5 library to support intelligent I/O buffering for deep memory and storage hierarchy systems☆34Feb 17, 2025Updated last year
- GNU Gzip with Kunpeng optimization.☆12Mar 30, 2022Updated 4 years ago
- A Deep Learning Meta-Framework and HPC Benchmarking Library☆81May 23, 2022Updated 4 years ago
- ☆24Nov 27, 2025Updated 8 months ago
- 📝 "Synthesizing Benchmarks for Predictive Modeling" (🥇 CGO'17 Best Paper)☆22Feb 10, 2023Updated 3 years ago
- Mirror site speedtest☆12Dec 4, 2023Updated 2 years ago
- ☆26Jan 10, 2023Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [READ ONLY] Refer to gitlab repo for updated version - Total Knowledge of I/O Reference Implementation. Please see wiki for contribution…☆22May 18, 2022Updated 4 years ago
- Domain-specific framework for performance analysis of parallel programs☆25Mar 23, 2026Updated 4 months ago
- ☆20Dec 8, 2022Updated 3 years ago
- A flexible C++ formatting library designed for i18n, using embedded script to output plural forms, grammatical gender, etc. correctly☆12Updated this week
- Used for testing the metadata performance of a file system☆26Nov 29, 2017Updated 8 years ago
- Linux kernel driver to export the TSC frequency via sysfs☆54Sep 24, 2019Updated 6 years ago
- PET: Optimizing Tensor Programs with Partially Equivalent Transformations and Automated Corrections☆126Jun 23, 2022Updated 4 years ago