High Performance Int8 GEMM Kernels for SM80 and later GPUs.
β24Mar 11, 2025Updated last year
Alternatives and similar repositories for gemm-int8
Users that are interested in gemm-int8 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- High Performance FP8 GEMM Kernels for SM89 and later GPUs.β22Jan 24, 2025Updated last year
- HALO: Hadamard-Assisted Low-Precision Optimization and Training method for finetuning LLMs. π The official implementation of https://arxβ¦β31Feb 17, 2025Updated last year
- This repository presents the source code for the paper "MILLION: Mastering Long-Context LLM Inference Via Outlier-Immunized KV Product Quβ¦β26Apr 2, 2025Updated last year
- Linear Attention for Efficient Bidirectional Sequence Modelingβ18May 13, 2025Updated last year
- This repository contains the official code for Energy Transformer---an efficient Energy-based Transformer variant for graph classificatioβ¦β27Jan 28, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- PyTorch Quantization Framework For OCP MX Datatypes.β16May 30, 2025Updated last year
- IntLLaMA: A fast and light quantization solution for LLaMAβ19Jul 21, 2023Updated 3 years ago
- Computer programming - ShanghaiTechβ12Jan 10, 2020Updated 6 years ago
- iEDA water-drop training initiativeβ13Sep 10, 2024Updated last year
- Pytorch implementation of "spectro-temporal attention-based voice activity detection"β13Jun 4, 2024Updated 2 years ago
- Transformer Architecture written with CUDA, C++ and LibTorch.β11Jul 26, 2025Updated last year
- GOMIL: Global Optimization of Multiplier by Integer Linear Programmingβ13Aug 25, 2021Updated 4 years ago
- ADAPTIVE RESONANCE THEORY. Gail A. Carpenter and Stephen Grossbergβ10Feb 10, 2015Updated 11 years ago
- Find, list, and inspect processes from Go (golang).β10Feb 4, 2018Updated 8 years ago
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Example of a full DC synthesis script for a simple designβ14Feb 25, 2019Updated 7 years ago
- [HPCA 2026] A GPU-optimized system for efficient long-context LLMs decoding with low-bit KV cache.β96May 14, 2026Updated 2 months ago
- A basic SAT solver implementation for the Logics in Informatics courseβ10Mar 8, 2015Updated 11 years ago
- Innovus backend scriptsβ16Jun 20, 2022Updated 4 years ago
- An extention to the GaLore paper, to perform Natural Gradient Descent in low rank subspaceβ19Oct 21, 2024Updated last year
- Faster Pytorch bitsandbytes 4bit fp4 nn.Linear opsβ30Mar 16, 2024Updated 2 years ago
- β12Apr 3, 2023Updated 3 years ago
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]β12Nov 8, 2024Updated last year
- Surgical GPU kernel benchmark: 7 hard problems, frontier coding agents, roofline-graded against hardware peak.β19Jun 12, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- β14May 4, 2026Updated 3 months ago
- β157Jun 22, 2023Updated 3 years ago
- CUDA 8-bit Tensor Core Matrix Multiplication based on m16n16k16 WMMA APIβ37Sep 15, 2023Updated 2 years ago
- Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.β15Aug 31, 2023Updated 2 years ago
- [ICML 2024] Code for the paper "Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy Biases"β38Jul 12, 2024Updated 2 years ago
- My solution code to parallel architecture and programming Spring 2016β12Aug 15, 2016Updated 9 years ago
- A DAG processor and compiler for a tree-based spatial datapath.β16Aug 24, 2022Updated 3 years ago
- β40Feb 28, 2020Updated 6 years ago
- LLM implementation one matrix multiplication at a timeβ13Aug 8, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Korean Morphological Analyzer using Hidden Markov Model (HMM)β10Nov 4, 2018Updated 7 years ago
- Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costsβ25Nov 11, 2025Updated 9 months ago
- β19Mar 10, 2022Updated 4 years ago
- Llama causal LM fully recreated in LibTorch. Designed to be used in Unreal Engine 5β16Sep 19, 2024Updated last year
- λ μ΄μ±κ²μ(μΉ΄νΈλΌμ΄λ) μμ¨μ£Όν μΈκ³΅μ§λ₯ κ°λ°β10Jun 30, 2022Updated 4 years ago
- Out-of-distribution Detection via Generation - NeurIPS 2019β18Oct 5, 2019Updated 6 years ago
- β19Nov 11, 2024Updated last year