Code appendix to an OpenCL matrix-multiplication tutorial
☆179Feb 7, 2017Updated 9 years ago
Alternatives and similar repositories for myGEMM
Users that are interested in myGEMM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Tuned OpenCL BLAS☆1,187Apr 13, 2026Updated 4 months ago
- CLTune: An automatic OpenCL & CUDA kernel tuner☆186Dec 12, 2022Updated 3 years ago
- a software library containing BLAS functions written in OpenCL☆866Aug 2, 2024Updated 2 years ago
- A portable high-level API with CUDA or OpenCL back-end☆56Oct 8, 2017Updated 8 years ago
- Sample program to compare calculation performance between CPU and GPU☆16Oct 27, 2016Updated 9 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Assembler for NVIDIA Maxwell architecture☆1,074Jan 3, 2023Updated 3 years ago
- AES-based random number generator in C☆11Apr 27, 2015Updated 11 years ago
- A synthetic micro-benchmark that measures peak compute, bandwidth, and matrix throughput of GPUs and CPUs☆511Updated this week
- The repository targets the OpenCL gemm function performance optimization. It compares several libraries clBLAS, clBLAST, MIOpenGemm, Inte…☆17Mar 28, 2019Updated 7 years ago
- Open single and half precision gemm implementations☆396Apr 2, 2023Updated 3 years ago
- OpenCL tool to detect buffer overflows in GPU kernels☆23Jan 7, 2019Updated 7 years ago
- Sequential and parallel GEMM implementations with C interface + Benchmark.☆12May 24, 2016Updated 10 years ago
- Learn OpenCL step by step.☆141Aug 30, 2022Updated 3 years ago
- ☆2,033Jul 29, 2023Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Caffe deep learning framework - optimized for Xeon Phi☆14May 12, 2015Updated 11 years ago
- assembler for NVIDIA FERMI. Imported from Google Code☆78Mar 22, 2015Updated 11 years ago
- The SHOC Benchmark Suite☆262Oct 6, 2025Updated 10 months ago
- A simple example of using the SDAccel build flow for AWS EC2's F1 instance type. Trys to avoid magic makefiles.☆10Aug 27, 2017Updated 8 years ago
- Winograd-based convolution implementation in OpenCL☆29Jan 22, 2017Updated 9 years ago
- This is a tuned sparse matrix dense vector multiplication(SpMV) library☆23Mar 21, 2016Updated 10 years ago
- BLAS OpenCL implementation.☆17Apr 8, 2015Updated 11 years ago
- ☆28Oct 26, 2019Updated 6 years ago
- New batched algorithm for sparse matrix-matrix multiplication (SpMM)☆16May 7, 2019Updated 7 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- OpenCL API, OpenCL C, Extensions, SPIR-V Environment Specs, Ref page, and C++ for OpenCL doc sources.☆416Aug 11, 2026Updated last week
- Mirror of http://gitlab.hpcrl.cse.ohio-state.edu/chong/ppopp19_ae, refactoring for understanding☆17Oct 20, 2021Updated 4 years ago
- Supplement of the ICFP'22 paper "‘do’ Unchained: Embracing Local Imperativity in a Purely Functional Language"☆16Feb 15, 2025Updated last year
- ☆32Aug 24, 2022Updated 3 years ago
- An implementation of SGEMV with performance comparable to cuBLAS.☆12May 21, 2021Updated 5 years ago
- load word embeddings to Torch.Tensor☆14May 12, 2016Updated 10 years ago
- Set of OpenCL microbenchmarks☆29Nov 19, 2025Updated 8 months ago
- Benchmark for Co-running Single Applications on Integrated Architectures☆12Jul 7, 2016Updated 10 years ago
- This repository contains notebooks showing how to perform mixed precision training in tf.keras 2.0☆12Dec 15, 2019Updated 6 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Cayley Dickson algebra implementation in python☆13Jan 3, 2019Updated 7 years ago
- Lecture Slide Issue Tracking☆258May 20, 2018Updated 8 years ago
- OpenCL memory benchmark☆15Dec 21, 2016Updated 9 years ago
- MAFIA: Multiple Application Framework for GPU architectures☆28Jan 21, 2022Updated 4 years ago
- a heterogeneous multiGPU level-3 BLAS library☆46Dec 9, 2019Updated 6 years ago
- An OpenCL device simulator and debugger☆373Mar 24, 2026Updated 4 months ago
- OpenCL Programming Examples☆22Jul 21, 2018Updated 8 years ago