code for benchmarking GPU performance based on cublasSgemm and cublasHgemm
☆35May 20, 2022Updated 4 years ago
Alternatives and similar repositories for cublasgemm-benchmark
Users that are interested in cublasgemm-benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- TLB Benchmarks☆35Sep 11, 2017Updated 9 years ago
- A GPU benchmark suite for assessing on-chip GPU memory bandwidth☆113Aug 12, 2017Updated 9 years ago
- A simple tool to profile performance of multiple combinations of GEMM of cuBLAS☆25Feb 9, 2021Updated 5 years ago
- A web app for monitoring Sun Grid Engine (SGE) cluster status☆12May 24, 2019Updated 7 years ago
- PHPQstat is a web interface that allows you to use the commands of the Sun Grid Engine (SGE) batch queue system (and forked projects). Wi…☆17Jan 17, 2017Updated 9 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- This example starts with a simple sum reduction in CUDA, then steps through a series of optimizations we can perform to improve its perfo…☆14Jun 8, 2020Updated 6 years ago
- To make it easy to benchmark AI accelerators☆192Dec 27, 2022Updated 3 years ago
- ☆15May 29, 2020Updated 6 years ago
- Tutorial on installing QEMU to simulate Zynq Devices with Petalinux☆24Jun 6, 2017Updated 9 years ago
- 为HSNW源码加上了详细的注释☆20Oct 26, 2022Updated 3 years ago
- ☆10Jun 9, 2017Updated 9 years ago
- Build a SystemVerilog Environment for an ALU, using OOP testbench components as; stimulus generator, driver, monitor, scoreboard. ALU was…☆10Mar 4, 2023Updated 3 years ago
- Image downsampler using a Lanczos filter implemented in ISPC☆15Sep 25, 2026Updated last week
- An auto-evolution framework that optimizes anything — your 7×24 team of algorithm engineers.☆1,040Sep 30, 2026Updated last week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Some source code about matrix multiplication implementation on CUDA☆34Sep 12, 2018Updated 8 years ago
- A comprehensive Visual Studio MSBuild integration of the Intel SPMD Compiler (ISPC), Premake support and a collection of ISPC tests and d…☆13Apr 7, 2025Updated last year
- Nonograms puzzle game written in Vala.☆12Dec 12, 2024Updated last year
- Benchmark code for the "Online normalizer calculation for softmax" paper☆113Jul 27, 2018Updated 8 years ago
- ☆10Jan 24, 2019Updated 7 years ago
- NGC Container Replicator☆30Dec 26, 2022Updated 3 years ago
- Samples and documentation for deploying EDA computing environments in AWS☆31Feb 6, 2026Updated 8 months ago
- Applications of the Teg differentiable programming language to problems spanning graphics and physical simulation.☆12Dec 30, 2021Updated 4 years ago
- Efficient and stable Determinant Quantum Monte Carlo simulations in Python☆11Sep 28, 2026Updated last week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Serverless setup using node.js☆14Jun 8, 2021Updated 5 years ago
- ☆10Jan 21, 2018Updated 8 years ago
- Energy-based Dropout and Pruning of Deep Neural Networks☆10Oct 9, 2020Updated 5 years ago
- A collection of papers about physically-based animation for deformable bodies☆13Feb 16, 2022Updated 4 years ago
- eBPF kernels and user space tools for BeagleBone SBCs☆10Jan 16, 2022Updated 4 years ago
- autonomous driving contest reference kit☆10Dec 2, 2021Updated 4 years ago
- ☆14Jan 14, 2020Updated 6 years ago
- Microbenchmarks for Aarch64 (Cortex A53)☆12Apr 19, 2023Updated 3 years ago
- Generate JSON and HTML system call table for aarch64 from Linux source.☆11Mar 6, 2021Updated 5 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A Benchmark Suite for Heterogeneous System Computation☆57Feb 20, 2025Updated last year
- Express DLA implementation for FPGA, revised based on NVDLA.☆11Oct 17, 2019Updated 6 years ago
- Examples illustrating usage of the rocBLAS library☆17Aug 12, 2024Updated 2 years ago
- ☆12Mar 13, 2023Updated 3 years ago
- Template for LaTeX beamer slides using #uulm corporate design.☆16Dec 3, 2022Updated 3 years ago
- A simple high performance CUDA GEMM implementation.☆441Jan 4, 2024Updated 2 years ago
- This is the repository containing the implementation of sparse dense matrix multiplication for the matrix dimension of 560 x 560.☆10Jul 7, 2021Updated 5 years ago