Source for Demystifying GPU Microarchitecture through Microbenchmarking
☆18May 29, 2023Updated 3 years ago
Alternatives and similar repositories for cudabmk
Users that are interested in cudabmk are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆25Jun 24, 2022Updated 4 years ago
- ☆26May 13, 2015Updated 11 years ago
- Benchmarks for locking algorithms as well as implementations of locking algorithms.☆26Mar 6, 2018Updated 8 years ago
- DeepPerf is a set of cuda assembling developing tools☆11Dec 19, 2018Updated 7 years ago
- A quick way to benchmark your CUDA compiler on a Linux environment☆27Mar 16, 2011Updated 15 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆10Mar 24, 2022Updated 4 years ago
- Sparse matrix computation library for GPU☆59Jul 12, 2020Updated 6 years ago
- Proof-of-concept implementation for the paper "Reviving Meltdown 3a" (ESORICS 2023)☆17Sep 25, 2023Updated 2 years ago
- Function hook, code injection, monitoring☆14Jul 12, 2018Updated 8 years ago
- ☆12May 9, 2024Updated 2 years ago
- Automatic Mapping Generation, Verification, and Exploration for ISA-based Spatial Accelerators☆125Oct 26, 2022Updated 3 years ago
- Repository holding the code base to AC-SpGEMM : "Adaptive Sparse Matrix-Matrix Multiplication on the GPU"☆32Jul 7, 2020Updated 6 years ago
- ☆24Mar 22, 2018Updated 8 years ago
- [ICLR 2021: Spotlight] Source code for the paper "A Panda? No, It's a Sloth: Slowdown Attacks on Adaptive Multi-Exit Neural Network Infer…☆15Feb 16, 2022Updated 4 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Accelerating DNN Convolutional Layers with Micro-batches☆63Apr 30, 2020Updated 6 years ago
- An EDM-enabled PHY + a rack-level network simulator☆14Dec 11, 2024Updated last year
- An implementation of the Latent Skill Embedding model☆10Feb 19, 2016Updated 10 years ago
- Source code of the U-TRR methodology presented in "Uncovering In-DRAM RowHammer Protection Mechanisms: A New Methodology, Custom RowHamme…☆19Nov 15, 2022Updated 3 years ago
- ☆12Oct 25, 2022Updated 3 years ago
- Eliminating Keystroke Timing Attacks☆22Dec 12, 2017Updated 8 years ago
- Linear-Time Self Attention with Codeword Histogram for Efficient Recommendation☆11Mar 23, 2021Updated 5 years ago
- reversing mtk-su☆17Mar 4, 2020Updated 6 years ago
- raccoon engine for kha [formerly lkl]☆12Sep 2, 2019Updated 7 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆16Aug 25, 2026Updated 3 weeks ago
- The translator that supports translating NVPTX to SPIR-V. This translator is modified from LLVM-SPIR-V Translator.☆45Oct 25, 2021Updated 4 years ago
- Library with JIT (Just-in-time) compilation support to optimize performance of small and medium matrix multiplication☆14Apr 27, 2021Updated 5 years ago
- A GPU performance prediction toolkit for CUDA programs☆18Mar 25, 2019Updated 7 years ago
- A SoC for DOOM☆20Apr 11, 2021Updated 5 years ago
- Superscalar Out-of-Order NPU Design on FPGA☆17May 17, 2024Updated 2 years ago
- Official Implementation of SEA: Sparse Linear Attention with Estimated Attention Mask (ICLR 2024)☆12Jun 20, 2025Updated last year
- Implemented a two-level (L1 and L2) cache simulator in C++ with round robin eviction policy☆10Jan 4, 2017Updated 9 years ago
- maxas Scott Grey's maxas assembler sgemm explaining the (for me) missing parts https://github.com/NervanaSystems/maxas☆17Dec 22, 2018Updated 7 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- TiledKernel is a code generation library based on macro kernels and memory hierarchy graph data structure.☆19May 12, 2024Updated 2 years ago
- Main repository for Harvard CS260r 2017.☆12Apr 25, 2017Updated 9 years ago
- A configurable general purpose graphics processing unit for☆12May 18, 2019Updated 7 years ago
- Yet another Game Boy emulator☆22Sep 4, 2022Updated 4 years ago
- Binsec/Haunted is an extension of Binsec to verify speculative constant-time and detect Spectre attacks.☆18Oct 19, 2023Updated 2 years ago
- ☆10May 12, 2022Updated 4 years ago
- Signal Flow Graph Solver in javascript☆16Aug 5, 2012Updated 14 years ago