FP64 equivalent GEMM by the Ozaki scheme with Int8 Tensor Cores
☆128Dec 2, 2025Updated 8 months ago
Alternatives and similar repositories for ozIMMU
Users that are interested in ozIMMU are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Acceleration codes for the Ozaki-scheme on integer matrix multiplication units.☆27Dec 10, 2025Updated 8 months ago
- GEMMul8 (GEMMulate): GEMM emulation and its extension to BLAS-like matrix operations using INT8/FP8 matrix engines based on the Ozaki Sch…☆88Jul 12, 2026Updated last month
- Fast SGEMM emulation on Tensor Cores☆17Feb 16, 2025Updated last year
- CLI-based multi-agents for Auto-Tuning (e.g. HPC code optimazation loops) supporting Local LLMs☆39Mar 25, 2026Updated 4 months ago
- This repository mirrors the principal Gitlab repository of the Chebyshev Accelerated Subspace iteration Eigensolver. If you want to contr…☆21Jul 8, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A gemm_tutorial☆38May 31, 2025Updated last year
- Fast implementation of memcpy☆12May 17, 2020Updated 6 years ago
- ☆36Mar 31, 2025Updated last year
- An extension library of WMMA API (Tensor Core API)☆115Jul 12, 2024Updated 2 years ago
- Benchmark your NCNN models on 3DS(or crash)☆10Apr 15, 2024Updated 2 years ago
- Handy tools & graphics API abstraction for blazing fast prototyping☆10Jan 17, 2024Updated 2 years ago
- ExBLAS: fast, accurate, and reproducible BLAS☆17Sep 13, 2021Updated 4 years ago
- Hierarchical Tensor Networks at Exascale☆69Jul 24, 2023Updated 3 years ago
- CUDA 12.2 HMM demos☆21Jul 26, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Usage of Eigen library with CMake.☆17Sep 18, 2018Updated 7 years ago
- ASM generation tool for GAS/NASM/MASM with Xbyak-like syntax in Python☆13Aug 3, 2026Updated last week
- Source-to-Source Debuggable Derivatives in Pure Python☆15Jan 23, 2024Updated 2 years ago
- variPEPS -- Versatile tensor network library for variational ground state simulations in two spatial dimensions☆19Updated this week
- GPU移植のための実装例(直接法に基づくN体計算)☆17Apr 11, 2026Updated 4 months ago
- Massively parallel tensor network solver☆59Updated this week
- Official repository for the paper "Exploring the Promise and Limits of Real-Time Recurrent Learning" (ICLR 2024)☆13Jun 11, 2025Updated last year
- Code for the paper "Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns"☆18Mar 15, 2024Updated 2 years ago
- ☆20May 30, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Distributed Communication-Optimal Shuffle and Transpose Algorithm☆14Apr 18, 2026Updated 3 months ago
- [OSDI' 26] Efficient LLM Serving on Commodity GPU Clusters with Data-Reduced Cross-Instance Orchestration☆24Jul 5, 2026Updated last month
- Benchmark tests supporting the TiledCUDA library.☆19Nov 19, 2024Updated last year
- Solving two-dimensional spin models with tensor networks (powered by PyTorch)☆105May 6, 2026Updated 3 months ago
- A simple Rust library implementing a 3D vector type☆11Dec 11, 2018Updated 7 years ago
- ☆18Dec 5, 2025Updated 8 months ago
- Embroid: Unsupervised Prediction Smoothing Can Improve Few-Shot Classification☆11Aug 12, 2023Updated 3 years ago
- A library for code transformations with guaranteed legality☆18Jun 12, 2026Updated 2 months ago
- automatic GPU offload for scientific libraries☆18Jun 17, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Curse-of-memory phenomenon of RNNs in sequence modelling☆19May 8, 2025Updated last year
- SLATE is a distributed, GPU-accelerated, dense linear algebra library targetting current and upcoming high-performance computing (HPC) sy…☆135Oct 21, 2025Updated 9 months ago
- Official code for UnICORNN (ICML 2021)☆28Oct 1, 2021Updated 4 years ago
- Daisytuner Optimizing Compiler Collection (docc)☆22Updated this week
- ☆23Aug 17, 2021Updated 4 years ago
- Sample codes for the PEPS excitation using generating function☆12Feb 16, 2026Updated 5 months ago
- train with kittens!☆67Oct 25, 2024Updated last year