High Performance Int8 GEMM Kernels for SM80 and later GPUs.
β24Mar 11, 2025Updated last year
Alternatives and similar repositories for gemm-int8
Users that are interested in gemm-int8 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- High Performance FP8 GEMM Kernels for SM89 and later GPUs.β21Jan 24, 2025Updated last year
- HALO: Hadamard-Assisted Low-Precision Optimization and Training method for finetuning LLMs. π The official implementation of https://arxβ¦β32Feb 17, 2025Updated last year
- Linear Attention for Efficient Bidirectional Sequence Modelingβ18May 13, 2025Updated last year
- PyTorch Quantization Framework For OCP MX Datatypes.β16May 30, 2025Updated last year
- This repository contains the official code for Energy Transformer---an efficient Energy-based Transformer variant for graph classificatioβ¦β28Jan 28, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A fork of the PEFT library, supporting Robust Adaptation (RoSA)β15Aug 16, 2024Updated 2 years ago
- IntLLaMA: A fast and light quantization solution for LLaMAβ19Jul 21, 2023Updated 3 years ago
- Computer programming - ShanghaiTechβ12Jan 10, 2020Updated 6 years ago
- iEDA water-drop training initiativeβ14Sep 10, 2024Updated 2 years ago
- A simple benchmark for modern image formats on mobiles, including WebP/HEIC/BPG/FLIF/AVIFβ12Oct 31, 2019Updated 6 years ago
- A double-double and quad-double package for Fortran and C++β26May 7, 2026Updated 4 months ago
- Pytorch implementation of "spectro-temporal attention-based voice activity detection"β13Jun 4, 2024Updated 2 years ago
- Transformer Architecture written with CUDA, C++ and LibTorch.β11Jul 26, 2025Updated last year
- β31Jun 15, 2022Updated 4 years ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Find, list, and inspect processes from Go (golang).β10Feb 4, 2018Updated 8 years ago
- Example of a full DC synthesis script for a simple designβ14Feb 25, 2019Updated 7 years ago
- [HPCA 2026] A GPU-optimized system for efficient long-context LLMs decoding with low-bit KV cache.β102May 14, 2026Updated 4 months ago
- A basic SAT solver implementation for the Logics in Informatics courseβ10Mar 8, 2015Updated 11 years ago
- Innovus backend scriptsβ16Jun 20, 2022Updated 4 years ago
- Faster Pytorch bitsandbytes 4bit fp4 nn.Linear opsβ30Mar 16, 2024Updated 2 years ago
- β12Apr 3, 2023Updated 3 years ago
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]β12Nov 8, 2024Updated last year
- Surgical GPU kernel benchmark: 7 hard problems, frontier coding agents, roofline-graded against hardware peak.β19Jun 12, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- β14May 4, 2026Updated 4 months ago
- β158Jun 22, 2023Updated 3 years ago
- CUDA 8-bit Tensor Core Matrix Multiplication based on m16n16k16 WMMA APIβ37Sep 15, 2023Updated 3 years ago
- Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.β15Aug 31, 2023Updated 3 years ago
- [ICML 2024] Code for the paper "Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy Biases"β38Jul 12, 2024Updated 2 years ago
- My solution code to parallel architecture and programming Spring 2016β12Aug 15, 2016Updated 10 years ago
- A DAG processor and compiler for a tree-based spatial datapath.β17Aug 24, 2022Updated 4 years ago
- LLM implementation one matrix multiplication at a timeβ14Aug 8, 2024Updated 2 years ago
- Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costsβ25Nov 11, 2025Updated 10 months ago
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Korean Morphological Analyzer using Hidden Markov Model (HMM)β10Nov 4, 2018Updated 7 years ago
- β19Mar 10, 2022Updated 4 years ago
- Llama causal LM fully recreated in LibTorch. Designed to be used in Unreal Engine 5β16Sep 19, 2024Updated 2 years ago
- λ μ΄μ±κ²μ(μΉ΄νΈλΌμ΄λ) μμ¨μ£Όν μΈκ³΅μ§λ₯ κ°λ°β10Jun 30, 2022Updated 4 years ago
- β89Apr 18, 2025Updated last year
- β19Nov 11, 2024Updated last year
- Out-of-distribution Detection via Generation - NeurIPS 2019β18Oct 5, 2019Updated 6 years ago