Row-wise block scaling for fp8 quantization matrix multiplication. Solution to GPU mode AMD challenge.
☆19Feb 9, 2026Updated 7 months ago
Alternatives and similar repositories for fp8-quant-matmul
Users that are interested in fp8-quant-matmul are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- My submission for the GPUMODE/AMD fp8 mm challenge☆29Jun 4, 2025Updated last year
- ☆19Jun 6, 2025Updated last year
- Efficient implementation of DeepSeek Ops (Blockwise FP8 GEMM, MoE, and MLA) for AMD Instinct MI300X☆80Feb 11, 2026Updated 7 months ago
- ☆14Dec 22, 2024Updated last year
- NVFP4 Flash-Attention 4 on BlackWell☆66Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- AMD-SHARK Inference Modeling and Serving☆71Jul 9, 2026Updated 2 months ago
- Write a fast kernel and see how you compare against the best humans and AI on gpumode.com☆114Sep 18, 2026Updated last week
- Personal solutions to the Triton Puzzles☆22Jul 18, 2024Updated 2 years ago
- ☆19Mar 29, 2026Updated 5 months ago
- TileFusion is an experimental C++ macro kernel template library that elevates the abstraction level in CUDA C for tile processing.☆119Aug 4, 2026Updated last month
- ☆76Updated this week
- ☆23Jun 12, 2025Updated last year
- ☆24Updated this week
- ☆15Oct 2, 2025Updated 11 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Tensara's GPU programming problems☆21Apr 23, 2026Updated 5 months ago
- [NeurIPS'25] Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning☆17Dec 12, 2025Updated 9 months ago
- ☆21Jun 12, 2026Updated 3 months ago
- ☆16Aug 5, 2025Updated last year
- Super fast FP32 matrix multiplication on RDNA3☆92Mar 30, 2025Updated last year
- A lightweight triton-based General Matrix Multiplication (GEMM) library.☆68Jul 21, 2026Updated 2 months ago
- ☆15Nov 19, 2025Updated 10 months ago
- ☆15Jan 27, 2025Updated last year
- Transforming Video Diffusion with Temporal Sparse Attention☆56Apr 8, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆13Sep 10, 2026Updated 2 weeks ago
- Official repo for "Binary Retrieval-augmented Reward Mitigates Hallucinations"☆16Nov 13, 2025Updated 10 months ago
- General Matrix Multiplication using NVIDIA Tensor Cores☆29Jan 25, 2025Updated last year
- Quantized LLM training in pure CUDA/C++.☆260Aug 31, 2026Updated 3 weeks ago
- ☆91Dec 16, 2025Updated 9 months ago
- Some funny cute/cuteDSL code snippets☆35Mar 2, 2026Updated 6 months ago
- GPU Functional Descriptor for memory access☆34May 24, 2026Updated 4 months ago
- ☆30Jan 7, 2026Updated 8 months ago
- [ICML 2025 Poster] SAE-V: Interpreting Multimodal Models for Enhanced Alignment☆18Jun 5, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Official Problem Sets / Reference Kernels for the GPU MODE Leaderboard!☆310Sep 18, 2026Updated last week
- ☆32Jul 2, 2025Updated last year
- Kolmogorov-Arnold Attention: Is Learnable Attention Better for Vision Transformers?☆16Jul 9, 2025Updated last year
- ☆46Oct 2, 2025Updated 11 months ago
- High-performance GPU kernels for LLM inference in OpenAI Triton. Fused RMSNorm, SwiGLU, INT8 GEMM with benchmarks and roofline analysis.☆43Jul 22, 2026Updated 2 months ago
- SMASH: Physics-guided Reconstruction of Collisions from Videos, SIGGRAPH Asia 2016☆11Jan 25, 2018Updated 8 years ago
- A curated reading list for researchers in the Philosophy of Interpretability☆19Aug 17, 2025Updated last year