Acceleration codes for the Ozaki-scheme on integer matrix multiplication units.
☆27Dec 10, 2025Updated 9 months ago
Alternatives and similar repositories for accelerator_for_ozIMMU
Users that are interested in accelerator_for_ozIMMU are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- GEMMul8 (GEMMulate): GEMM emulation and its extension to BLAS-like matrix operations using INT8/FP8 matrix engines based on the Ozaki Sch…☆92Updated this week
- Fast SGEMM emulation on Tensor Cores☆17Feb 16, 2025Updated last year
- ☆37Mar 31, 2025Updated last year
- [OSDI' 26] Efficient LLM Serving on Commodity GPU Clusters with Data-Reduced Cross-Instance Orchestration☆25Jul 5, 2026Updated 2 months ago
- automatic GPU offload for scientific libraries☆18Jun 17, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Towards a million-node RISC-V cluster.☆14Mar 6, 2025Updated last year
- FFVC - Frontflow/violet Cartesian☆14Apr 5, 2020Updated 6 years ago
- [CF ’20] Verified Instruction-Level Energy Consumption Measurement for NVIDIA GPUs☆15Dec 11, 2020Updated 5 years ago
- GPULZ: Optimizing LZSS Lossless Compression for Multi-byte Data on Modern GPUs☆16Apr 18, 2025Updated last year
- OpenACC for Python☆20Jul 17, 2019Updated 7 years ago
- ☆22Jul 8, 2024Updated 2 years ago
- ☆17Dec 5, 2024Updated last year
- Cosmic Tagging Network for Neutrino Physics☆13Aug 22, 2026Updated 3 weeks ago
- Itoyori: A distributed multi-threading runtime system for global-view fork-join task parallelism☆23Feb 9, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [CVPR 2023] "TrojViT: Trojan Insertion in Vision Transformers" by Mengxin Zheng, Qian Lou, Lei Jiang☆15Jan 5, 2024Updated 2 years ago
- 上海交通大学Xflops超算队2024招新第一轮考核试题☆16Oct 15, 2024Updated last year
- A generic implementation of tensor einsum in Fortran.☆30Jan 29, 2021Updated 5 years ago
- Panopticon is a complete in-DRAM RowHammer mitigation. This code simulates an implementation of Panopticon in DDR5.☆14Jun 2, 2023Updated 3 years ago
- ☆27Jun 29, 2026Updated 2 months ago
- ☆17Apr 9, 2025Updated last year
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]☆12Nov 8, 2024Updated last year
- Tutorials for Timemory☆21Aug 1, 2024Updated 2 years ago
- Efficient Auto-scalable Scientific Infrastructure for Engineers and Researchers☆16Jul 10, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- variPEPS -- Versatile tensor network library for variational ground state simulations in two spatial dimensions☆20Sep 9, 2026Updated last week
- ☆31Aug 31, 2026Updated 2 weeks ago
- Bluespec environment for working with the ulx3s board and its lattice ecp5 fpga☆15Mar 9, 2025Updated last year
- A DAG processor and compiler for a tree-based spatial datapath.☆17Aug 24, 2022Updated 4 years ago
- ☆53Apr 27, 2026Updated 4 months ago
- A runtime-independent crate for transforming Wasm-DWARF☆12Mar 11, 2020Updated 6 years ago
- ☆29Nov 15, 2025Updated 10 months ago
- TransPimLib is a library for transcendental (and other hard-to-calculate) functions in general-purpose PIM systems, TransPimLib provides …☆15Apr 21, 2023Updated 3 years ago
- ☆29Sep 11, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆28Oct 11, 2022Updated 3 years ago
- [NSDI25] AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN Training☆36May 2, 2025Updated last year
- PTX on XPUs☆135Jun 15, 2026Updated 3 months ago
- ucas hpc course code☆15May 24, 2023Updated 3 years ago
- Enhancing the convergence speed by 2x and improving the training success of Physics-Informed Neural Networks (PINNs).☆14Oct 14, 2024Updated last year
- Light-weight real-time multi-object detection and tracking in Nvidia TX2☆10May 10, 2019Updated 7 years ago
- MY BLOG☆16Updated this week