A CUDA kernel optimization toolkit for validation, benchmarking, Nsight Compute profiling, bottleneck analysis, and iterative tuning. It helps improve custom GPU operators with reproducible workflows and evidence-based performance comparison.
☆203Sep 5, 2026Updated this week
Alternatives and similar repositories for cuda-optimized-skill
Users that are interested in cuda-optimized-skill are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Skills for writing tilelang and debugging with CUDA toolkits.☆137May 20, 2026Updated 3 months ago
- ☆270Updated this week
- ☆165Aug 8, 2026Updated 3 weeks ago
- PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.☆228Dec 24, 2025Updated 8 months ago
- NVSHMEM‑Tutorial: Build a DeepEP‑like GPU Buffer☆200Feb 11, 2026Updated 6 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Agentic Kernel Optimization for All — automated GPU kernel optimization for any kernel, any hardware, any language☆362Aug 17, 2026Updated 3 weeks ago
- ☆441Aug 26, 2026Updated last week
- ☆787Aug 23, 2026Updated 2 weeks ago
- High Performance LLM Inference Operator Library☆1,146Updated this week
- Automated CUDA kernel performance diagnostics from NVIDIA Nsight Compute (NCU) CSV exports.☆35Mar 18, 2026Updated 5 months ago
- ☆92Feb 5, 2026Updated 7 months ago
- A skill for automatically optimizing CUDA code.☆43Mar 26, 2026Updated 5 months ago
- ☆955Updated this week
- ☆220Aug 26, 2026Updated last week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Distributed Compiler and Optimized Parallel Kernels☆1,539Aug 12, 2026Updated 3 weeks ago
- FLA but cuTile☆27Apr 17, 2026Updated 4 months ago
- From Minimal GEMM to Everything☆241Jul 9, 2026Updated last month
- A benchmark of real-world DL kernel problems☆291Jul 15, 2026Updated last month
- NVIDIA cuTile learn☆168Dec 9, 2025Updated 8 months ago
- fake CUTLASS to get peformance☆26Apr 28, 2026Updated 4 months ago
- Benchmark tests supporting the TiledCUDA library.☆19Nov 19, 2024Updated last year
- let coding agents use ncu skills analysis cuda program automatically!☆125May 25, 2026Updated 3 months ago
- ☆32Jul 2, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Low overhead tracing library and trace visualizer for pipelined CUDA kernels☆137Jul 29, 2026Updated last month
- ☆100Mar 31, 2026Updated 5 months ago
- Tile-Based Runtime for Ultra-Low-Latency LLM Inference☆1,756Aug 13, 2026Updated 3 weeks ago
- mKernel: fast multi-node, multi-GPU fused kernels☆270Updated this week
- FlashKDA: high-performance Kimi Delta Attention kernels☆1,244Updated this week
- Winner 🏆 (Agent-only) MLSys 2026 - FlashInfer AI Kernel Generation Contest for the DeepSeek Sparse Attention (DSA) track with an average…☆159Aug 21, 2026Updated 2 weeks ago
- Autonomous GPU kernel optimization system driven by AI agents.☆31Mar 29, 2026Updated 5 months ago
- High-performance LLM operator library built on TileLang.☆178Updated this week
- Building the Virtuous Cycle for AI-driven LLM Systems☆281May 1, 2026Updated 4 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation☆1,311Jul 8, 2026Updated last month
- Implement Flash Attention using Cute.☆111Dec 17, 2024Updated last year
- A kernel library written in tilelang☆1,753Apr 23, 2026Updated 4 months ago
- Accelerating MoE with IO and Tile-aware Optimizations☆759Aug 29, 2026Updated last week
- ☆48Nov 1, 2025Updated 10 months ago
- Examples of CUDA implementations by Cutlass CuTe☆282Jul 1, 2025Updated last year
- LLM Inference via Triton (Flexible & Modular): Focused on Kernel Optimization using CUBIN binaries, Starting from gpt-oss Model☆119Apr 28, 2026Updated 4 months ago