Build compute kernels and load them from the Hub.
β715Jul 20, 2026Updated this week
Alternatives and similar repositories for kernels
Users that are interested in kernels are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π· Build compute kernelsβ213Apr 6, 2026Updated 3 months ago
- Kernel sources for https://huggingface.co/kernels-communityβ129Updated this week
- A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.β910Updated this week
- A Quirky Assortment of CuTe Kernelsβ1,063Updated this week
- FlashInfer: Kernel Library for LLM Servingβ5,988Updated this week
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Minimalistic large language model 3D-parallelism trainingβ2,755May 26, 2026Updated last month
- Accelerating MoE with IO and Tile-aware Optimizationsβ732Jul 4, 2026Updated 2 weeks ago
- Efficient Triton Kernels for LLM Trainingβ6,528Updated this week
- Tile primitives for speedy kernelsβ3,552Jul 13, 2026Updated last week
- Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.β361Updated this week
- π Efficient implementations for emerging model architecturesβ5,379Updated this week
- PyTorch native quantization and sparsity for training and inferenceβ2,909Updated this week
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)β1,148Mar 24, 2026Updated 3 months ago
- Mirage Persistent Kernel: Compiling LLMs into a MegaKernelβ2,376Updated this week
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Minimalistic 4D-parallelism distributed training framework for education purposeβ2,254Aug 26, 2025Updated 10 months ago
- Kernels, of the mega variety :)β780May 26, 2026Updated last month
- A PyTorch native platform for training generative AI modelsβ5,545Updated this week
- Helpful tools and examples for working with flex-attentionβ1,209Updated this week
- π yet another mixture of expertsβ23Jun 5, 2026Updated last month
- Hugging Face Jobsβ20Jul 11, 2025Updated last year
- Distributed Compiler based on Triton for Parallel Systemsβ1,494Updated this week
- Mount Hugging Face Buckets and repos as local filesystems. No download, no copy, no waiting.β767Jul 12, 2026Updated last week
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernelsβ6,674Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- A lightweight, local-first, and free experiment tracking library from Hugging Face π€β1,584Updated this week
- Autonomous GPU Kernel Generation & Optimization via Deep Agentsβ486Updated this week
- Helpful kernel tutorials, examples and SKILLs for tile-based GPU programmingβ776Updated this week
- FlashKDA: high-performance Kimi Delta Attention kernelsβ462May 26, 2026Updated last month
- A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hβ¦β3,435Updated this week
- Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backendsβ2,486Jun 29, 2026Updated 3 weeks ago
- CUDA Templates and Python DSLs for High-Performance Linear Algebraβ10,104Updated this week
- MSLK (Meta Superintelligence Labs Kernels) is a collection of PyTorch GPU operator libraries that are designed and optimized for GenAI trβ¦β121Updated this week
- torchcomms: a modern PyTorch communications APIβ377Updated this week
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Development repository for the Triton language and compilerβ19,738Updated this week
- π Accelerate inference and training of π€ Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimizationβ¦β3,448Updated this week
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.β1,469Mar 19, 2026Updated 4 months ago
- FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.β1,109Sep 4, 2024Updated last year
- Fast and memory-efficient exact attentionβ24,497Updated this week
- Fast low-bit matmul kernels in Tritonβ477Updated this week
- β28May 26, 2026Updated last month