Matrix Multiplication on GPU using Shared Memory considering Coalescing and Bank Conflicts
☆26Aug 29, 2022Updated 3 years ago
Alternatives and similar repositories for Cuda-Matrix-Multiplication
Users that are interested in Cuda-Matrix-Multiplication are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This is a c++ implementation of an LSTM Neural Network parallelized for a GPU using CUDA☆25Oct 29, 2017Updated 8 years ago
- A collection of awesome algorithms, implemented in CUDA.☆26Feb 6, 2018Updated 8 years ago
- ☆14Nov 3, 2025Updated 8 months ago
- An expression template based linear algebra library running completely on the GPU using CUDA☆26Jun 24, 2021Updated 5 years ago
- Parallel selection on GPUs☆15Mar 23, 2021Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Implementation of lid driven cavity solver based on SIMPLE algorithm☆15Jan 11, 2019Updated 7 years ago
- Matrix-Vector Multiplication Using Shared and Coalesced Memory Access☆16Apr 9, 2013Updated 13 years ago
- Open source skill library for AI coding agents to write, optimize, and debug high performance compute kernels across CUDA, Triton, and qu…☆47Jun 21, 2026Updated last month
- 本仓库在OpenVINO推理框架下部署Nanodet检测算法,并重写预处理和后处理部分,具有超高性能!让你在Intel CPU平台上的检测速度起飞! 并基于NNCF和PPQ工具将模型量化(PTQ)至int8精度,推理速度更快!☆16Jun 14, 2023Updated 3 years ago
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆24Jan 30, 2026Updated 5 months ago
- CUDA C simple application for Nvidia's GPU☆11Jun 7, 2022Updated 4 years ago
- A new QR decomposition algorithm implemented in CUDA☆18Jun 24, 2024Updated 2 years ago
- Simple and efficient memory pool is implemented with C++11.☆10Jun 2, 2022Updated 4 years ago
- Repository holding the code base to AC-SpGEMM : "Adaptive Sparse Matrix-Matrix Multiplication on the GPU"☆31Jul 7, 2020Updated 6 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- OCCA Python API: JIT Compilation for Multiple Architectures☆11Dec 20, 2019Updated 6 years ago
- SGEMM and DGEMM subroutines using AVX512F instructions.☆15May 22, 2022Updated 4 years ago
- Matlab mex wrappers to cuSPARSE (NVIDIA)☆11Dec 10, 2025Updated 7 months ago
- Musings in GEMM (General Matrix Multiplication)☆14Dec 14, 2025Updated 7 months ago
- Record GPU memory accesses of a CUDA program and visualize the access pattern in a browser☆13Nov 17, 2020Updated 5 years ago
- Clink is a library that provides APIs and infrastructure to facilitate the development of parallelizable feature engineering operators th…☆30Feb 21, 2022Updated 4 years ago
- GEMV implementation with CUTLASS☆21Aug 21, 2025Updated 11 months ago
- A tutorial/example of the Python C-API and integration with CUDA kernels.☆14Jul 7, 2019Updated 7 years ago
- Matrix multiplication on GPUs for matrices stored on a CPU. Similar to cublasXt, but ported to both NVIDIA and AMD GPUs.☆33Apr 2, 2025Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- TensorRT-FastSAM(https://github.com/CASIA-IVA-Lab/FastSAM)☆23Feb 29, 2024Updated 2 years ago
- 对 tensorRT_Pro 开源项目理解☆22Feb 23, 2023Updated 3 years ago
- Personal Notes for Learning HPC & Parallel Computation [NO LONGER ADDING NEW CONTENT]☆78Jul 29, 2022Updated 3 years ago
- ☆34Jul 23, 2024Updated 2 years ago
- Source code of the IPDPS '21 paper: "TileSpMV: A Tiled Algorithm for Sparse Matrix-Vector Multiplication on GPUs" by Yuyao Niu, Zhengyang…☆13Aug 12, 2022Updated 3 years ago
- ☆10Jul 4, 2022Updated 4 years ago
- It's a project combined with hardware and software, the goal is to make a smart watch based on esp8266 chip. The smart watch has so many …☆10Jul 9, 2019Updated 7 years ago
- Android Face Recognition uses Microsoft Project Oxford Face API for face detection and identification.☆12Nov 13, 2015Updated 10 years ago
- This project provides example FeatHub (https://github.com/alibaba/feathub) programs☆28Sep 21, 2023Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- This is a repository to practice multi-thread programming in C++☆31Feb 21, 2024Updated 2 years ago
- In this project, both dynamic and thermal simulation of laser powder bed fusion are implemented using CUDA.☆24Dec 28, 2025Updated 6 months ago
- MIMO precoding and detection algorithm☆19Feb 26, 2018Updated 8 years ago
- Dark channel Haze removal algorithm with CUDA acceleration (typically 10x or more speedup using a Nvidia GPU)☆14Dec 7, 2017Updated 8 years ago
- Implementation of Inverse Propensity Matrix Factorization with Pytorch-Lightning☆12Sep 23, 2020Updated 5 years ago
- lshash for python3☆10Mar 21, 2018Updated 8 years ago
- nVidia's CUDA accelerated Spin Transformations of Discrete Surfaces, based on the original code and paper by Keenan Crane, Ulrich Pinkall…☆17Mar 14, 2018Updated 8 years ago