High Performance Int8 GEMM Kernels for SM80 and later GPUs.
β23Mar 11, 2025Updated last year
Alternatives and similar repositories for gemm-int8
Users that are interested in gemm-int8 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- High Performance FP8 GEMM Kernels for SM89 and later GPUs.β21Jan 24, 2025Updated last year
- HALO: Hadamard-Assisted Low-Precision Optimization and Training method for finetuning LLMs. π The official implementation of https://arxβ¦β31Feb 17, 2025Updated last year
- This repository presents the source code for the paper "MILLION: Mastering Long-Context LLM Inference Via Outlier-Immunized KV Product Quβ¦β25Apr 2, 2025Updated last year
- This repository contains the official code for Energy Transformer---an efficient Energy-based Transformer variant for graph classificatioβ¦β27Jan 28, 2024Updated 2 years ago
- PyTorch Quantization Framework For OCP MX Datatypes.β16May 30, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- IntLLaMA: A fast and light quantization solution for LLaMAβ19Jul 21, 2023Updated 3 years ago
- Computer programming - ShanghaiTechβ12Jan 10, 2020Updated 6 years ago
- iEDA water-drop training initiativeβ14Sep 10, 2024Updated last year
- A simple benchmark for modern image formats on mobiles, including WebP/HEIC/BPG/FLIF/AVIFβ12Oct 31, 2019Updated 6 years ago
- A double-double and quad-double package for Fortran and C++β23May 7, 2026Updated 2 months ago
- Transformer Architecture written with CUDA, C++ and LibTorch.β11Jul 26, 2025Updated 11 months ago
- β31Jun 15, 2022Updated 4 years ago
- GOMIL: Global Optimization of Multiplier by Integer Linear Programmingβ13Aug 25, 2021Updated 4 years ago
- ADAPTIVE RESONANCE THEORY. Gail A. Carpenter and Stephen Grossbergβ10Feb 10, 2015Updated 11 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Find, list, and inspect processes from Go (golang).β10Feb 4, 2018Updated 8 years ago
- [HPCA 2026] A GPU-optimized system for efficient long-context LLMs decoding with low-bit KV cache.β96May 14, 2026Updated 2 months ago
- A basic SAT solver implementation for the Logics in Informatics courseβ10Mar 8, 2015Updated 11 years ago
- Innovus backend scriptsβ15Jun 20, 2022Updated 4 years ago
- An extention to the GaLore paper, to perform Natural Gradient Descent in low rank subspaceβ19Oct 21, 2024Updated last year
- Faster Pytorch bitsandbytes 4bit fp4 nn.Linear opsβ30Mar 16, 2024Updated 2 years ago
- β12Apr 3, 2023Updated 3 years ago
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]β12Nov 8, 2024Updated last year
- Visualize machine learning models with Netron in VSCodeβ19Apr 22, 2026Updated 2 months ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- β13May 4, 2026Updated 2 months ago
- β157Jun 22, 2023Updated 3 years ago
- CUDA 8-bit Tensor Core Matrix Multiplication based on m16n16k16 WMMA APIβ37Sep 15, 2023Updated 2 years ago
- β19Sep 5, 2024Updated last year
- Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.β15Aug 31, 2023Updated 2 years ago
- [ICML 2024] Code for the paper "Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy Biases"β38Jul 12, 2024Updated 2 years ago
- My solution code to parallel architecture and programming Spring 2016β12Aug 15, 2016Updated 9 years ago
- A DAG processor and compiler for a tree-based spatial datapath.β16Aug 24, 2022Updated 3 years ago
- β40Feb 28, 2020Updated 6 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Korean Morphological Analyzer using Hidden Markov Model (HMM)β10Nov 4, 2018Updated 7 years ago
- LLM implementation one matrix multiplication at a timeβ13Aug 8, 2024Updated last year
- Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costsβ25Nov 11, 2025Updated 8 months ago
- β19Mar 10, 2022Updated 4 years ago
- λ μ΄μ±κ²μ(μΉ΄νΈλΌμ΄λ) μμ¨μ£Όν μΈκ³΅μ§λ₯ κ°λ°β10Jun 30, 2022Updated 4 years ago
- Llama causal LM fully recreated in LibTorch. Designed to be used in Unreal Engine 5β16Sep 19, 2024Updated last year
- β89Apr 18, 2025Updated last year