Public skills collected from well-known open-source projects focused on LLM infrastructure, GPU kernels, compiler/operator development
☆31May 7, 2026Updated 4 months ago
Alternatives and similar repositories for Awesome-Kernel-Skills
Users that are interested in Awesome-Kernel-Skills are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆23Jun 29, 2026Updated 2 months ago
- ☆15Apr 28, 2026Updated 4 months ago
- See vLLM official support: https://github.com/vllm-project/vllm-ascend☆11Feb 5, 2025Updated last year
- ☆13Jul 15, 2024Updated 2 years ago
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆555Sep 8, 2026Updated last week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- SGLang is a high-performance serving framework for large language models and multimodal models.☆22Updated this week
- Source code and datasets of "Efficient GPU-Accelerated Subgraph Matching", accepted by SIGMOD'23 - By Xibo Sun and Prof. Qiong Luo☆20Jul 20, 2023Updated 3 years ago
- Source codes of "Fast Continuous Subgraph Matching over Streaming Graphs via Backtracking Reduction", SIGMOD 2023☆14Sep 7, 2023Updated 3 years ago
- Adaptive Topology Reconstruction for Robust Graph Representation Learning [Efficient ML Model]☆10Feb 11, 2025Updated last year
- 开个坑,啥时候有时间啥时候写☆13Oct 26, 2023Updated 2 years ago
- Adamas: Hadamard Sparse Attention for Efficient Long-context Inference☆15May 19, 2026Updated 4 months ago
- papers about recommender system.☆10May 18, 2021Updated 5 years ago
- Triton language and compiler for Ascend NPU☆166Updated this week
- The official implementation of NOSA☆19Jun 11, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- TurboQuant reference implementation — KV cache compression with engineering insights (ICLR 2026 paper reproduction)☆17Mar 28, 2026Updated 5 months ago
- A web overlay mod for MU3 which exports some real-time gameplay information☆25Nov 3, 2025Updated 10 months ago
- MultiArchKernelBench: A Multi-Platform Benchmark for Kernel Generation☆70Jul 8, 2026Updated 2 months ago
- Python xml.sax☆10Mar 19, 2019Updated 7 years ago
- ☆29Feb 2, 2017Updated 9 years ago
- let coding agents use ncu skills analysis cuda program automatically!☆127May 25, 2026Updated 3 months ago
- Trust: Triangle Counting Reloaded on GPUs☆21Oct 14, 2023Updated 2 years ago
- Official implementation of "DPad: Efficient Diffusion Language Models with Suffix Dropout"☆66Feb 13, 2026Updated 7 months ago
- A "standard library" of Triton kernels.☆26Oct 2, 2025Updated 11 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [AAAI 2026] Official implementation of "FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models". If you find this reposi…☆19Aug 26, 2026Updated 3 weeks ago
- ☆73Mar 24, 2026Updated 5 months ago
- FA4-based Relative Attention Kernel developed by TML and Colfax☆18Sep 11, 2026Updated last week
- ☆21Jul 22, 2022Updated 4 years ago
- Triton Compiler related materials.☆46Mar 16, 2026Updated 6 months ago
- Review automated kernel generation in the era of LLMs☆310Jun 25, 2026Updated 2 months ago
- A framework and CLI toolkit for orchestrating teams of loosely-coupled AI agents.☆19Aug 9, 2026Updated last month
- 安卓版 V2Ray 客户端,支持 Xray core 和 v2fly core☆20May 9, 2024Updated 2 years ago
- Surgical GPU kernel benchmark: 7 hard problems, frontier coding agents, roofline-graded against hardware peak.☆19Jun 12, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Multi-Level Triton Runner supporting Python, IR, PTX, AMDGCN, cubin and hasco.☆100Updated this week
- 《Linux 高性能服务器编程》 书中实例☆17Dec 29, 2016Updated 9 years ago
- Official Repo for the VideoVerse☆15Mar 29, 2026Updated 5 months ago
- [KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML…☆210Mar 29, 2026Updated 5 months ago
- 基于SG2300X的视频检索【使用自然语言搜索视频内容,定位到符合描述的具体时间段】☆13Feb 29, 2024Updated 2 years ago
- Flash-Linear-Attention models beyond language☆21Aug 28, 2025Updated last year
- 适用于sophon bm1684x,基于 Langchain 与 ChatGLM 等语言模型的本地知识库问答☆13Jun 5, 2024Updated 2 years ago