Acceleration codes for the Ozaki-scheme on integer matrix multiplication units.
☆27Dec 10, 2025Updated 9 months ago
Alternatives and similar repositories for accelerator_for_ozIMMU
Users that are interested in accelerator_for_ozIMMU are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- GEMMul8 (GEMMulate): GEMM emulation and its extension to BLAS-like matrix operations using INT8/FP8 matrix engines based on the Ozaki Sch…☆102Updated this week
- FP64 equivalent GEMM by the Ozaki scheme with Int8 Tensor Cores☆130Dec 2, 2025Updated 10 months ago
- Fast SGEMM emulation on Tensor Cores☆17Feb 16, 2025Updated last year
- ☆37Mar 31, 2025Updated last year
- [OSDI' 26] Efficient LLM Serving on Commodity GPU Clusters with Data-Reduced Cross-Instance Orchestration☆27Jul 5, 2026Updated 3 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- The quasiparticle self-consistent GW method in the PMT method (LAPW+LMTO+Lo).☆37Updated this week
- Fast implementation of memcpy☆12May 17, 2020Updated 6 years ago
- automatic GPU offload for scientific libraries☆19Jun 17, 2026Updated 3 months ago
- Towards a million-node RISC-V cluster.☆14Mar 6, 2025Updated last year
- A double-double and quad-double package for Fortran and C++☆28Updated this week
- High Availability Shared Pipeline Engine☆17Sep 15, 2023Updated 3 years ago
- GPULZ: Optimizing LZSS Lossless Compression for Multi-byte Data on Modern GPUs☆16Apr 18, 2025Updated last year
- OpenACC for Python☆20Jul 17, 2019Updated 7 years ago
- ☆22Jul 8, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆17Dec 5, 2024Updated last year
- 2023 XFlops Training☆13Jan 23, 2024Updated 2 years ago
- Cosmic Tagging Network for Neutrino Physics☆13Aug 22, 2026Updated last month
- Itoyori: A distributed multi-threading runtime system for global-view fork-join task parallelism☆23Feb 9, 2024Updated 2 years ago
- The LLVM Project is a collection of modular and reusable compiler and toolchain technologies. Note: the repository does not accept github…☆34Jul 20, 2021Updated 5 years ago
- [CVPR 2023] "TrojViT: Trojan Insertion in Vision Transformers" by Mengxin Zheng, Qian Lou, Lei Jiang☆15Jan 5, 2024Updated 2 years ago
- A generic implementation of tensor einsum in Fortran.☆30Jan 29, 2021Updated 5 years ago
- Panopticon is a complete in-DRAM RowHammer mitigation. This code simulates an implementation of Panopticon in DDR5.☆15Jun 2, 2023Updated 3 years ago
- ☆17Apr 9, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]☆12Nov 8, 2024Updated last year
- Tutorials for Timemory☆21Aug 1, 2024Updated 2 years ago
- Efficient Auto-scalable Scientific Infrastructure for Engineers and Researchers☆16Jul 10, 2026Updated 2 months ago
- variPEPS -- Versatile tensor network library for variational ground state simulations in two spatial dimensions☆20Updated this week
- Bluespec environment for working with the ulx3s board and its lattice ecp5 fpga☆15Mar 9, 2025Updated last year
- A DAG processor and compiler for a tree-based spatial datapath.☆17Aug 24, 2022Updated 4 years ago
- Examples showing how to utilize the NVML library for GPU monitoring☆29May 31, 2022Updated 4 years ago
- A software implementation of an O-RAN split-7.2 compatible Radio Unit using software-defined Radios (SDRs).☆34Aug 16, 2026Updated last month
- ☆56Apr 27, 2026Updated 5 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆64May 4, 2024Updated 2 years ago
- TransPimLib is a library for transcendental (and other hard-to-calculate) functions in general-purpose PIM systems, TransPimLib provides …☆15Apr 21, 2023Updated 3 years ago
- Communication-Avoiding Recursive Matrix Multiply☆19Jul 10, 2013Updated 13 years ago
- ☆28Oct 11, 2022Updated 3 years ago
- ☆13Jul 27, 2026Updated 2 months ago
- [NSDI25] AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN Training☆36May 2, 2025Updated last year
- PTX on XPUs☆138Jun 15, 2026Updated 3 months ago