Official Repo of CudaForge
☆87Dec 2, 2025Updated 8 months ago
Alternatives and similar repositories for CudaForge
Users that are interested in CudaForge are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆25Feb 14, 2026Updated 6 months ago
- GraphDancer: Training LLMs to Explore and Reason over Graphs via Curriculum Reinforcement Learning☆20May 25, 2026Updated 3 months ago
- LLM4Kernel: A Survey of Large Language Models for GPU Kernel Development☆82Mar 31, 2026Updated 5 months ago
- ☆102Nov 22, 2025Updated 9 months ago
- Skills for writing tilelang and debugging with CUDA toolkits.☆137May 20, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆532Jul 15, 2026Updated last month
- CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning☆320Nov 3, 2025Updated 9 months ago
- A benchmark of real-world DL kernel problems☆288Jul 15, 2026Updated last month
- ☆162Aug 8, 2026Updated 3 weeks ago
- ☆12Aug 4, 2022Updated 4 years ago
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)☆1,221Mar 24, 2026Updated 5 months ago
- Building the Virtuous Cycle for AI-driven LLM Systems☆277May 1, 2026Updated 3 months ago
- ☆208Updated this week
- ☆16Jan 24, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Some funny cute/cuteDSL code snippets☆34Mar 2, 2026Updated 5 months ago
- Speed of Light Analysis for ML Model Runtime☆114Aug 22, 2026Updated last week
- [KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML…☆205Mar 29, 2026Updated 5 months ago
- ☆414Updated this week
- This is code for How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and Analysis☆18Nov 5, 2025Updated 9 months ago
- PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.☆228Dec 24, 2025Updated 8 months ago
- Automated High-Performance GPU Kernel Generation☆136Jun 1, 2026Updated 2 months ago
- ☆19May 9, 2025Updated last year
- Evaluating Large Language Models for CUDA Code Generation ComputeEval is a framework designed to generate and evaluate CUDA code from Lar…☆147Aug 13, 2026Updated 2 weeks ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆98Mar 31, 2026Updated 5 months ago
- ScienceMeter: Tracking Scientific Knowledge Updates in Language Models, COLM 2026☆17Jun 28, 2025Updated last year
- CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation☆1,264Jul 8, 2026Updated last month
- Optimize tensor program fast with Felix, a gradient descent autotuner.☆33Mar 5, 2026Updated 5 months ago
- ☆155Aug 18, 2025Updated last year
- Expert Specialization MoE Solution based on CUTLASS☆27Apr 14, 2026Updated 4 months ago
- A reference implementation of the Mind Mappings Framework.☆30Dec 2, 2021Updated 4 years ago
- Automated CUDA kernel performance diagnostics from NVIDIA Nsight Compute (NCU) CSV exports.☆35Mar 18, 2026Updated 5 months ago
- This repository provides the official implementation of QSVD, a method for efficient low-rank approximation that unifies Query-Key-Value …☆28May 16, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [MLSys 26] 🥇 Solution for Gated Delta Net Track of MLSys 26 Flash infer competition☆36May 22, 2026Updated 3 months ago
- ☆92Feb 5, 2026Updated 6 months ago
- [MLSys 2022] "BNS-GCN: Efficient Full-Graph Training of Graph Convolutional Networks with Partition-Parallelism and Random Boundary Node …☆56Oct 6, 2023Updated 2 years ago
- Ship correct and fast LLM kernels to PyTorch☆154Jan 14, 2026Updated 7 months ago
- ☆257Updated this week
- Sample Codes using NVSHMEM on Multi-GPU☆30Jan 22, 2023Updated 3 years ago
- Benchmarking Open-Ended Inference Optimization by AI Agents☆43Jul 6, 2026Updated last month