High Performance Int8 GEMM Kernels for SM80 and later GPUs.
โ24Mar 11, 2025Updated last year
Alternatives and similar repositories for gemm-int8
Users that are interested in gemm-int8 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- High Performance FP8 GEMM Kernels for SM89 and later GPUs.โ21Jan 24, 2025Updated last year
- HALO: Hadamard-Assisted Low-Precision Optimization and Training method for finetuning LLMs. ๐ The official implementation of https://arxโฆโ32Feb 17, 2025Updated last year
- This repository presents the source code for the paper "MILLION: Mastering Long-Context LLM Inference Via Outlier-Immunized KV Product Quโฆโ26Updated this week
- This repository contains the official code for Energy Transformer---an efficient Energy-based Transformer variant for graph classificatioโฆโ27Jan 28, 2024Updated 2 years ago
- PyTorch Quantization Framework For OCP MX Datatypes.โ16May 30, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer โข AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A fork of the PEFT library, supporting Robust Adaptation (RoSA)โ15Aug 16, 2024Updated 2 years ago
- IntLLaMA: A fast and light quantization solution for LLaMAโ19Jul 21, 2023Updated 3 years ago
- iEDA water-drop training initiativeโ13Sep 10, 2024Updated last year
- A double-double and quad-double package for Fortran and C++โ26May 7, 2026Updated 3 months ago
- โ31Jun 15, 2022Updated 4 years ago
- GOMIL: Global Optimization of Multiplier by Integer Linear Programmingโ13Aug 25, 2021Updated 5 years ago
- Find, list, and inspect processes from Go (golang).โ10Feb 4, 2018Updated 8 years ago
- Example of a full DC synthesis script for a simple designโ14Feb 25, 2019Updated 7 years ago
- [HPCA 2026] A GPU-optimized system for efficient long-context LLMs decoding with low-bit KV cache.โ99May 14, 2026Updated 3 months ago
- Simple, predictable pricing with DigitalOcean hosting โข AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- A basic SAT solver implementation for the Logics in Informatics courseโ10Mar 8, 2015Updated 11 years ago
- An extention to the GaLore paper, to perform Natural Gradient Descent in low rank subspaceโ19Oct 21, 2024Updated last year
- Faster Pytorch bitsandbytes 4bit fp4 nn.Linear opsโ30Mar 16, 2024Updated 2 years ago
- โ12Apr 3, 2023Updated 3 years ago
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]โ12Nov 8, 2024Updated last year
- Surgical GPU kernel benchmark: 7 hard problems, frontier coding agents, roofline-graded against hardware peak.โ19Jun 12, 2026Updated 2 months ago
- โ157Jun 22, 2023Updated 3 years ago
- [ACL'24, Outstanding Paper] Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!โ39Aug 2, 2024Updated 2 years ago
- CUDA 8-bit Tensor Core Matrix Multiplication based on m16n16k16 WMMA APIโ37Sep 15, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient โข AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- โ19Sep 5, 2024Updated last year
- โ11Jul 20, 2023Updated 3 years ago
- Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.โ15Aug 31, 2023Updated 3 years ago
- [ICML 2024] Code for the paper "Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy Biases"โ38Jul 12, 2024Updated 2 years ago
- My solution code to parallel architecture and programming Spring 2016โ12Aug 15, 2016Updated 10 years ago
- โ40Feb 28, 2020Updated 6 years ago
- LLM implementation one matrix multiplication at a timeโ14Aug 8, 2024Updated 2 years ago
- Korean Morphological Analyzer using Hidden Markov Model (HMM)โ10Nov 4, 2018Updated 7 years ago
- Llama causal LM fully recreated in LibTorch. Designed to be used in Unreal Engine 5โ16Sep 19, 2024Updated last year
- Managed Kubernetes at scale on DigitalOcean โข AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ๋ ์ด์ฑ๊ฒ์(์นดํธ๋ผ์ด๋) ์์จ์ฃผํ ์ธ๊ณต์ง๋ฅ ๊ฐ๋ฐโ10Jun 30, 2022Updated 4 years ago
- โ89Apr 18, 2025Updated last year
- Out-of-distribution Detection via Generation - NeurIPS 2019โ18Oct 5, 2019Updated 6 years ago
- โ19Nov 11, 2024Updated last year
- Official code of "StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs".โ77Jun 23, 2025Updated last year
- ๅญๆพไธไบ CUDA ็ผ็จ็ธๅ ณ็ๅๅฎขๆไปถใโ22Oct 16, 2025Updated 10 months ago
- [NeurIPS 2024] For paper Parameter Competition Balancing for Model Mergingโ48Oct 11, 2024Updated last year