Build compute kernels and load them from the Hub.
β728Aug 28, 2026Updated this week
Alternatives and similar repositories for kernels
Users that are interested in kernels are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π· Build compute kernelsβ213Apr 6, 2026Updated 4 months ago
- Kernel sources for https://huggingface.co/kernels-communityβ141Updated this week
- A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.β929Updated this week
- A Quirky Assortment of CuTe Kernelsβ1,134Aug 21, 2026Updated last week
- Minimalistic large language model 3D-parallelism trainingβ2,804May 26, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- FlashInfer: Kernel Library for LLM Servingβ6,281Updated this week
- Tile primitives for speedy kernelsβ3,658Updated this week
- Hugging Face Jobsβ20Jul 11, 2025Updated last year
- Accelerating MoE with IO and Tile-aware Optimizationsβ750Updated this week
- Efficient Triton Kernels for LLM Trainingβ6,592Updated this week
- Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.β366Updated this week
- A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hβ¦β3,509Updated this week
- Helpful tools and examples for working with flex-attentionβ1,229Updated this week
- π Efficient implementations for emerging model architecturesβ5,659Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- PyTorch native quantization for training and inferenceβ2,959Updated this week
- Minimalistic 4D-parallelism distributed training framework for education purposeβ2,291Aug 26, 2025Updated last year
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)β1,220Mar 24, 2026Updated 5 months ago
- Mirage Persistent Kernel: Compiling LLMs into a MegaKernelβ2,461Updated this week
- Kernels, of the mega variety :)β814May 26, 2026Updated 3 months ago
- A PyTorch native platform for training generative AI modelsβ5,678Updated this week
- π yet another mixture of expertsβ23Jun 5, 2026Updated 2 months ago
- Distributed Compiler and Optimized Parallel Kernelsβ1,529Aug 12, 2026Updated 2 weeks ago
- Mount Hugging Face Buckets and repos as local filesystems. No download, no copy, no waiting.β789Aug 16, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernelsβ7,302Updated this week
- Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backendsβ2,530Aug 11, 2026Updated 2 weeks ago
- ANE accelerated embedding models!β20Dec 11, 2024Updated last year
- A lightweight, local-first, and free experiment tracking library from Hugging Face π€β1,666Updated this week
- Autonomous GPU Kernel Generation & Optimization via Deep Agentsβ532Jul 15, 2026Updated last month
- Helpful kernel tutorials, examples and SKILLs for tile-based GPU programmingβ806Updated this week
- FlashKDA: high-performance Kimi Delta Attention kernelsβ1,236Jul 30, 2026Updated last month
- CUDA Templates and Python DSLs for High-Performance Linear Algebraβ10,345Updated this week
- MSLK (Meta Superintelligence Labs Kernels) is a collection of PyTorch GPU operator libraries that are designed and optimized for GenAI trβ¦β149Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- β354Updated this week
- torchcomms: a modern PyTorch communications APIβ390Updated this week
- Development repository for the Triton language and compilerβ20,037Updated this week
- π Accelerate inference and training of π€ Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimizationβ¦β3,473Updated this week
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.β1,538Mar 19, 2026Updated 5 months ago
- FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.β1,136Sep 4, 2024Updated last year
- Fast and memory-efficient exact attentionβ24,801Updated this week