Build compute kernels and load them from the Hub.
β723Aug 8, 2026Updated this week
Alternatives and similar repositories for kernels
Users that are interested in kernels are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π· Build compute kernelsβ213Apr 6, 2026Updated 4 months ago
- Kernel sources for https://huggingface.co/kernels-communityβ136Updated this week
- A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.β918Updated this week
- A Quirky Assortment of CuTe Kernelsβ1,096Updated this week
- Minimalistic large language model 3D-parallelism trainingβ2,779May 26, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- FlashInfer: Kernel Library for LLM Servingβ6,133Updated this week
- Tile primitives for speedy kernelsβ3,621Jul 13, 2026Updated 3 weeks ago
- Hugging Face Jobsβ20Jul 11, 2025Updated last year
- Accelerating MoE with IO and Tile-aware Optimizationsβ737Jul 4, 2026Updated last month
- Efficient Triton Kernels for LLM Trainingβ6,558Updated this week
- Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.β364Updated this week
- A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hβ¦β3,484Updated this week
- Helpful tools and examples for working with flex-attentionβ1,222Updated this week
- π Efficient implementations for emerging model architecturesβ5,528Updated this week
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- PyTorch native quantization and sparsity for training and inferenceβ2,933Updated this week
- Minimalistic 4D-parallelism distributed training framework for education purposeβ2,274Aug 26, 2025Updated 11 months ago
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)β1,192Mar 24, 2026Updated 4 months ago
- Mirage Persistent Kernel: Compiling LLMs into a MegaKernelβ2,414Updated this week
- Kernels, of the mega variety :)β797May 26, 2026Updated 2 months ago
- A PyTorch native platform for training generative AI modelsβ5,604Updated this week
- π yet another mixture of expertsβ23Jun 5, 2026Updated 2 months ago
- Distributed Compiler based on Triton for Parallel Systemsβ1,512Updated this week
- Mount Hugging Face Buckets and repos as local filesystems. No download, no copy, no waiting.β777Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernelsβ7,167Updated this week
- Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backendsβ2,510Jun 29, 2026Updated last month
- ANE accelerated embedding models!β20Dec 11, 2024Updated last year
- A lightweight, local-first, and free experiment tracking library from Hugging Face π€β1,632Updated this week
- Autonomous GPU Kernel Generation & Optimization via Deep Agentsβ504Jul 15, 2026Updated 3 weeks ago
- Helpful kernel tutorials, examples and SKILLs for tile-based GPU programmingβ796Updated this week
- FlashKDA: high-performance Kimi Delta Attention kernelsβ1,193Jul 30, 2026Updated last week
- CUDA Templates and Python DSLs for High-Performance Linear Algebraβ10,217Updated this week
- MSLK (Meta Superintelligence Labs Kernels) is a collection of PyTorch GPU operator libraries that are designed and optimized for GenAI trβ¦β145Updated this week
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- β353Updated this week
- torchcomms: a modern PyTorch communications APIβ385Updated this week
- Development repository for the Triton language and compilerβ19,908Updated this week
- π Accelerate inference and training of π€ Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimizationβ¦β3,455Updated this week
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.β1,506Mar 19, 2026Updated 4 months ago
- FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.β1,123Sep 4, 2024Updated last year
- Fast and memory-efficient exact attentionβ24,658Updated this week