FP64 equivalent GEMM by the Ozaki scheme with Int8 Tensor Cores
☆127Dec 2, 2025Updated 9 months ago
Alternatives and similar repositories for ozIMMU
Users that are interested in ozIMMU are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Acceleration codes for the Ozaki-scheme on integer matrix multiplication units.☆27Dec 10, 2025Updated 8 months ago
- GEMMul8 (GEMMulate): GEMM emulation and its extension to BLAS-like matrix operations using INT8/FP8 matrix engines based on the Ozaki Sch…☆89Aug 13, 2026Updated 3 weeks ago
- Fast SGEMM emulation on Tensor Cores☆17Feb 16, 2025Updated last year
- A python package for Grassmann tensor network computation☆19Dec 12, 2024Updated last year
- A gemm_tutorial☆39May 31, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Fast implementation of memcpy☆12May 17, 2020Updated 6 years ago
- ☆37Mar 31, 2025Updated last year
- An extension library of WMMA API (Tensor Core API)☆115Jul 12, 2024Updated 2 years ago
- Benchmark your NCNN models on 3DS(or crash)☆10Apr 15, 2024Updated 2 years ago
- ExBLAS: fast, accurate, and reproducible BLAS☆17Sep 13, 2021Updated 4 years ago
- Hierarchical Tensor Networks at Exascale☆69Jul 24, 2023Updated 3 years ago
- Usage of Eigen library with CMake.☆17Sep 18, 2018Updated 7 years ago
- Source-to-Source Debuggable Derivatives in Pure Python☆15Jan 23, 2024Updated 2 years ago
- Massively parallel tensor network solver☆59Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official repository for the paper "Exploring the Promise and Limits of Real-Time Recurrent Learning" (ICLR 2024)☆13Jun 11, 2025Updated last year
- Code for the paper "Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns"☆18Mar 15, 2024Updated 2 years ago
- ☆20May 30, 2024Updated 2 years ago
- Distributed Communication-Optimal Shuffle and Transpose Algorithm☆14Updated this week
- [OSDI' 26] Efficient LLM Serving on Commodity GPU Clusters with Data-Reduced Cross-Instance Orchestration☆24Jul 5, 2026Updated last month
- Benchmark tests supporting the TiledCUDA library.☆19Nov 19, 2024Updated last year
- Yet another symmetric tensor network☆51Aug 26, 2026Updated last week
- A simple Rust library implementing a 3D vector type☆11Dec 11, 2018Updated 7 years ago
- ☆18Aug 19, 2026Updated 2 weeks ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A library for code transformations with guaranteed legality☆18Updated this week
- automatic GPU offload for scientific libraries☆18Jun 17, 2026Updated 2 months ago
- SLATE is a distributed, GPU-accelerated, dense linear algebra library targetting current and upcoming high-performance computing (HPC) sy…☆134Aug 27, 2026Updated last week
- oneAPI Deep Neural Network Library (oneDNN)☆10Feb 2, 2022Updated 4 years ago
- Call ncnn from Fortran☆18Dec 18, 2022Updated 3 years ago
- Official code for UnICORNN (ICML 2021)☆28Oct 1, 2021Updated 4 years ago
- Daisytuner Optimizing Compiler Collection (docc)☆22Updated this week
- ☆23Aug 17, 2021Updated 5 years ago
- Sample codes for the PEPS excitation using generating function☆14Feb 16, 2026Updated 6 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- train with kittens!☆67Oct 25, 2024Updated last year
- A Deeplearn Model to rec table in photo with ncnn. 一个深度学习模型用于检测图片中的表格 画像内のテーブルを検出するためのディープラーニング モデル☆20Mar 2, 2025Updated last year
- Software/library for simulations of quantum gates☆22Jul 27, 2026Updated last month
- ☆19Updated this week
- 🎃 GPU load-balancing library for regular and irregular computations.☆67Jun 25, 2026Updated 2 months ago
- 量子ソフトウェア産学協働ゼミ☆17Feb 12, 2026Updated 6 months ago
- Handwritten GEMM using Intel AMX (Advanced Matrix Extension)☆17Jan 11, 2025Updated last year