将 NVIDIA PTX ISA 9.1、CUDA 13.1 (Runtime/Driver)、Math API 13.x、cuBLAS 13.2 及 NCCL 官方文档转换为易于检索的 Markdown 格式,并提供配套的 AI IDE 技能库(支持 Claude Code、Trae 等),专为 GPU 底层开发与大语言模型 (LLM) 推理优化提供知识增强的自动化辅助。
☆24Jul 24, 2026Updated last month
Alternatives and similar repositories for cuda-code-skill
Users that are interested in cuda-code-skill are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A skill for automatically optimizing CUDA code.☆43Mar 26, 2026Updated 4 months ago
- A simple script to plot the Roofline model for given HW platforms and applications☆10Mar 17, 2026Updated 5 months ago
- ☆161Aug 8, 2026Updated 2 weeks ago
- PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.☆227Dec 24, 2025Updated 7 months ago
- SpInfer: Leveraging Low-Level Sparsity for Efficient Large Language Model Inference on GPUs☆70Mar 25, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- GEMV implementation with CUTLASS☆21Aug 21, 2025Updated last year
- Compression primitives for uplink compression in Federated Learning that are compatible with Secure Aggregation.☆11Jul 27, 2022Updated 4 years ago
- Agent application/benchmark/workload traces should be placed here.☆15Apr 13, 2026Updated 4 months ago
- Awesome code, projects, books, etc. related to CUDA☆39Aug 9, 2026Updated 2 weeks ago
- This is a repository to practice multi-thread programming in C++☆32Feb 21, 2024Updated 2 years ago
- let coding agents use ncu skills analysis cuda program automatically!☆120May 25, 2026Updated 2 months ago
- 存放一些 CUDA 编程相关的博客文件。☆22Oct 16, 2025Updated 10 months ago
- ☆16Sep 2, 2020Updated 5 years ago
- 复现 nanovllm并增加注释☆15Jan 27, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- TensorRT encapsulation, learn, rewrite, practice.☆31Oct 19, 2022Updated 3 years ago
- PyTorch extension enabling direct access to cuDNN-accelerated C++ convolution functions.☆13Mar 14, 2021Updated 5 years ago
- A tool for model sparse based on torch.fx☆13Jun 3, 2024Updated 2 years ago
- A direct convolution library targeting ARM multi-core CPUs.☆12Nov 27, 2024Updated last year
- Created a simple neural network using C++17 standard and the Eigen library that supports both forward and backward propagation.☆11Jul 27, 2024Updated 2 years ago
- Simple and efficient memory pool is implemented with C++11.☆10Jun 2, 2022Updated 4 years ago
- 在PyTorch上重构multi-agent deep deterministic policy gradient(MADDPG),将https://github.com/xuemei-ye/maddpg-mpe 修改到自己电脑上可运行。因为本人笔记本没有CUDA,实验速度…☆14May 10, 2019Updated 7 years ago
- pytorch版基于gpt+nezha的中文多轮Cdial☆11Oct 22, 2022Updated 3 years ago
- 数据库内核笔记☆14Aug 18, 2022Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Proxy based on QUIC.☆11Feb 3, 2022Updated 4 years ago
- Benchmark tests supporting the TiledCUDA library.☆19Nov 19, 2024Updated last year
- DeepSeek-V4 Lecture☆27Aug 10, 2026Updated 2 weeks ago
- GEMM☆10Aug 26, 2023Updated 2 years ago
- ☆23Aug 20, 2025Updated last year
- Flash Attention in ~100 lines of CUDA (forward pass only)☆12Jun 10, 2024Updated 2 years ago
- A practical way of learning Swizzle☆45Feb 3, 2025Updated last year
- ☆11May 2, 2023Updated 3 years ago
- An example implementatation of synchronized queue for inter-process communication in shared memory☆13Feb 17, 2017Updated 9 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A CUDA kernel optimization toolkit for validation, benchmarking, Nsight Compute profiling, bottleneck analysis, and iterative tuning. It …☆198Apr 22, 2026Updated 4 months ago
- A Chinese-focused PyTorch framework for exploring Attention Residuals in Qwen3-style causal LMs, with baseline, Block AttnRes, Full AttnR…☆21May 3, 2026Updated 3 months ago
- 方便扩展的Cuda算子理解和优化框架,仅用在学习使用☆18Jun 13, 2024Updated 2 years ago
- TensorRT☆11Sep 22, 2020Updated 5 years ago
- Some examples for bonsai_term☆24Jul 10, 2026Updated last month
- ☆11May 16, 2026Updated 3 months ago
- An onnx-based quantitation tool.☆71Jan 8, 2024Updated 2 years ago