Build compute kernels and load them from the Hub.
β745Sep 17, 2026Updated this week
Alternatives and similar repositories for kernels
Users that are interested in kernels are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π· Build compute kernelsβ213Apr 6, 2026Updated 5 months ago
- Kernel sources for https://huggingface.co/kernels-communityβ149Updated this week
- A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.β946Updated this week
- A Quirky Assortment of CuTe Kernelsβ1,146Updated this week
- Minimalistic large language model 3D-parallelism trainingβ2,825Updated this week
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- FlashInfer: Kernel Library for LLM Servingβ6,443Updated this week
- Tile primitives for speedy kernelsβ3,716Updated this week
- Hugging Face Jobsβ20Jul 11, 2025Updated last year
- Accelerating MoE with IO and Tile-aware Optimizationsβ769Aug 29, 2026Updated 3 weeks ago
- Efficient Triton Kernels for LLM Trainingβ6,618Updated this week
- Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.β369Updated this week
- A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hβ¦β3,540Updated this week
- Helpful tools and examples for working with flex-attentionβ1,247Updated this week
- π Efficient implementations for emerging model architecturesβ5,764Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- PyTorch native quantization for training and inferenceβ2,980Updated this week
- Minimalistic 4D-parallelism distributed training framework for education purposeβ2,305Aug 26, 2025Updated last year
- Mirage Persistent Kernel: Compiling LLMs into a MegaKernelβ2,503Sep 12, 2026Updated last week
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)β1,248Mar 24, 2026Updated 5 months ago
- Kernels, of the mega variety :)β829May 26, 2026Updated 3 months ago
- A PyTorch native platform for training generative AI modelsβ5,746Updated this week
- π yet another mixture of expertsβ23Jun 5, 2026Updated 3 months ago
- Distributed Compiler and Optimized Parallel Kernelsβ1,546Updated this week
- Mount Hugging Face Buckets and repos as local filesystems. No download, no copy, no waiting.β803Updated this week
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernelsβ7,431Updated this week
- Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backendsβ2,545Updated this week
- ANE accelerated embedding models!β20Dec 11, 2024Updated last year
- A lightweight, local-first, and free experiment tracking library from Hugging Face π€β1,685Updated this week
- Autonomous GPU Kernel Generation & Optimization via Deep Agentsβ555Sep 8, 2026Updated last week
- Helpful kernel tutorials, examples and SKILLs for tile-based GPU programmingβ810Updated this week
- CUDA Templates and Python DSLs for High-Performance Linear Algebraβ10,453Updated this week
- FlashKDA: high-performance Kimi Delta Attention kernelsβ1,254Sep 1, 2026Updated 2 weeks ago
- MSLK (Meta Superintelligence Labs Kernels) is a collection of PyTorch GPU operator libraries that are designed and optimized for GenAI trβ¦β154Updated this week
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- β355Updated this week
- torchcomms: a modern PyTorch communications APIβ396Updated this week
- Development repository for the Triton language and compilerβ20,190Updated this week
- π Accelerate inference and training of π€ Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimizationβ¦β3,492Updated this week
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.β1,560Mar 19, 2026Updated 6 months ago
- FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.β1,148Sep 4, 2024Updated 2 years ago
- Fast and memory-efficient exact attentionβ24,966Updated this week