Build compute kernels and load them from the Hub.
β764Oct 8, 2026Updated this week
Alternatives and similar repositories for kernels
Users that are interested in kernels are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π· Build compute kernelsβ213Apr 6, 2026Updated 6 months ago
- Kernel sources for https://huggingface.co/kernels-communityβ149Updated this week
- A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.β960Updated this week
- A Quirky Assortment of CuTe Kernelsβ1,166Sep 17, 2026Updated 3 weeks ago
- Minimalistic large language model 3D-parallelism trainingβ2,834Updated this week
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- FlashInfer: Kernel Library for LLM Servingβ6,556Updated this week
- Tile primitives for speedy kernelsβ3,746Sep 12, 2026Updated 3 weeks ago
- Hugging Face Jobsβ20Jul 11, 2025Updated last year
- Accelerating MoE with IO and Tile-aware Optimizationsβ777Aug 29, 2026Updated last month
- Efficient Triton Kernels for LLM Trainingβ6,650Updated this week
- Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.β370Oct 2, 2026Updated last week
- A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hβ¦β3,571Updated this week
- Helpful tools and examples for working with flex-attentionβ1,254Updated this week
- π Efficient implementations for emerging model architecturesβ5,833Updated this week
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- PyTorch native quantization for training and inferenceβ2,995Updated this week
- Minimalistic 4D-parallelism distributed training framework for education purposeβ2,320Aug 26, 2025Updated last year
- Mirage Persistent Kernel: Compiling LLMs into a MegaKernelβ2,542Updated this week
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)β1,281Mar 24, 2026Updated 6 months ago
- Kernels, of the mega variety :)β843May 26, 2026Updated 4 months ago
- A PyTorch native platform for training generative AI modelsβ5,788Updated this week
- π yet another mixture of expertsβ23Jun 5, 2026Updated 4 months ago
- Distributed Compiler and Optimized Parallel Kernelsβ1,559Sep 18, 2026Updated 3 weeks ago
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernelsβ8,516Updated this week
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Mount Hugging Face Buckets and repos as local filesystems. No download, no copy, no waiting.β812Updated this week
- Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backendsβ2,554Updated this week
- ANE accelerated embedding models!β20Dec 11, 2024Updated last year
- A lightweight, local-first, and free experiment tracking library from Hugging Face π€β1,713Updated this week
- Helpful kernel tutorials, examples and SKILLs for tile-based GPU programmingβ823Updated this week
- Autonomous GPU Kernel Generation & Optimization via Deep Agentsβ579Sep 8, 2026Updated last month
- CUDA Templates and Python DSLs for High-Performance Linear Algebraβ10,546Sep 23, 2026Updated 2 weeks ago
- FlashKDA: high-performance Kimi Delta Attention kernelsβ1,272Sep 1, 2026Updated last month
- MSLK (Meta Superintelligence Labs Kernels) is a collection of PyTorch GPU operator libraries that are designed and optimized for GenAI trβ¦β155Updated this week
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- β360Updated this week
- torchcomms: a modern PyTorch communications APIβ400Updated this week
- Development repository for the Triton language and compilerβ20,321Updated this week
- π Accelerate inference and training of π€ Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimizationβ¦β3,498Updated this week
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.β1,580Mar 19, 2026Updated 6 months ago
- FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.β1,153Sep 4, 2024Updated 2 years ago
- Fast and memory-efficient exact attentionβ25,106Updated this week