GEMMul8 (GEMMulate): GEMM emulation and its extension to BLAS-like matrix operations using INT8/FP8 matrix engines based on the Ozaki Scheme II
☆88Jul 12, 2026Updated 3 weeks ago
Alternatives and similar repositories for GEMMul8
Users that are interested in GEMMul8 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Acceleration codes for the Ozaki-scheme on integer matrix multiplication units.☆27Dec 10, 2025Updated 8 months ago
- Fast SGEMM emulation on Tensor Cores☆17Feb 16, 2025Updated last year
- The Hardware Sampling (hws) library can be used to track hardware performance like clock frequency, memory usage, temperatures, or power …☆21Jun 25, 2026Updated last month
- Itoyori: A distributed multi-threading runtime system for global-view fork-join task parallelism☆23Feb 9, 2024Updated 2 years ago
- ☆26Apr 13, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- OpenMP offload playground☆10Nov 16, 2024Updated last year
- ☆19Jun 26, 2026Updated last month
- Layered prefill changes the scheduling axis from tokens to layers and removes redundant MoE weight reloads while keeping decode stall fre…☆20Mar 9, 2026Updated 5 months ago
- ☆36Mar 31, 2025Updated last year
- An extension library of WMMA API (Tensor Core API)☆115Jul 12, 2024Updated 2 years ago
- A gemm_tutorial☆38May 31, 2025Updated last year
- ☆17Dec 5, 2024Updated last year
- The quasiparticle self-consistent GW method in the PMT method (LAPW+LMTO+Lo).☆36Jun 7, 2026Updated 2 months ago
- This repository mirrors the principal Gitlab repository of the Chebyshev Accelerated Subspace iteration Eigensolver. If you want to contr…☆21Jul 8, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Concurrent hash tries for C++ 14 with no memory management whatsoever.☆10Aug 30, 2016Updated 9 years ago
- A Heterogeneous GPU Platform for AI and Neural Graphics☆65Jun 22, 2026Updated last month
- Computing FLOPs with Intel Software Development Emulator (Intel SDE)☆27Oct 22, 2023Updated 2 years ago
- Communication-Avoiding Recursive Matrix Multiply☆19Jul 10, 2013Updated 13 years ago
- Repository with examples and exercises for OLCF and AMD's HIP training series☆17Oct 16, 2023Updated 2 years ago
- Exchange correlation (XC) library for density functional theory (DFT) calculations in modern C++☆28Jun 12, 2026Updated last month
- ☆48Apr 27, 2026Updated 3 months ago
- A place to store information for the tensor discussions and possible specifications.☆24Jun 17, 2026Updated last month
- Code generator for simint vectorized integrals☆29Mar 16, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A standalone implementation of the MPI Fortran 2018 module☆39Jun 21, 2026Updated last month
- PolyMage is a domain-specific language and optimizing code generator for auto-parallelisation☆14Jul 15, 2016Updated 10 years ago
- Chunky Loop Analyzer: A Polyhedral Representation Extraction Tool for High Level Programs☆26Dec 19, 2022Updated 3 years ago
- ☆11Jun 29, 2021Updated 5 years ago
- Anatomy of High-Performance GEMM with Online Fault Tolerance on GPUs☆14Apr 3, 2025Updated last year
- Simple "Stable Fluids" Implementation☆16Mar 25, 2024Updated 2 years ago
- An example Hardware Processing Engine☆12Feb 4, 2023Updated 3 years ago
- Quantum Transport Simulations at the Exascale and Beyond☆20Updated this week
- Flux tutorial slides and materials☆25Jul 21, 2026Updated 2 weeks ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- instruction-bench☆35Jan 10, 2023Updated 3 years ago
- A hierarchical collective communications library with portable optimizations☆38Dec 8, 2024Updated last year
- Lisp dialect designed for HPC and AI☆26Jul 3, 2026Updated last month
- Fast GPU based tensor core reductions☆12Jan 13, 2023Updated 3 years ago
- ☆116Updated this week
- ☆63May 4, 2024Updated 2 years ago
- ☆37Jun 11, 2026Updated last month