Official Repo of CudaForge
☆90Dec 2, 2025Updated 10 months ago
Alternatives and similar repositories for CudaForge
Users that are interested in CudaForge are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆27Feb 14, 2026Updated 7 months ago
- GraphDancer: Training LLMs to Explore and Reason over Graphs via Curriculum Reinforcement Learning☆21May 25, 2026Updated 4 months ago
- LLM4Kernel: A Survey of Large Language Models for GPU Kernel Development☆82Mar 31, 2026Updated 6 months ago
- ☆107Nov 22, 2025Updated 10 months ago
- Skills for writing tilelang and debugging with CUDA toolkits.☆145Sep 15, 2026Updated 3 weeks ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆580Sep 8, 2026Updated last month
- CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning☆324Nov 3, 2025Updated 11 months ago
- A benchmark of real-world DL kernel problems☆308Jul 15, 2026Updated 2 months ago
- ☆12Aug 4, 2022Updated 4 years ago
- ☆170Aug 8, 2026Updated 2 months ago
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)☆1,283Mar 24, 2026Updated 6 months ago
- Building the Virtuous Cycle for AI-driven LLM Systems☆289Sep 25, 2026Updated 2 weeks ago
- ☆16Jan 24, 2024Updated 2 years ago
- Some funny cute/cuteDSL code snippets☆35Mar 2, 2026Updated 7 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆249Aug 26, 2026Updated last month
- Speed of Light Analysis for ML Model Runtime☆131Updated this week
- [KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML…☆211Mar 29, 2026Updated 6 months ago
- This is code for How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and Analysis☆18Updated this week
- ☆501Aug 26, 2026Updated last month
- PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.☆229Dec 24, 2025Updated 9 months ago
- ☆19May 9, 2025Updated last year
- Automated High-Performance GPU Kernel Generation☆139Jun 1, 2026Updated 4 months ago
- Evaluating Large Language Models for CUDA Code Generation ComputeEval is a framework designed to generate and evaluate CUDA code from Lar…☆149Sep 14, 2026Updated 3 weeks ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆102Mar 31, 2026Updated 6 months ago
- ScienceMeter: Tracking Scientific Knowledge Updates in Language Models, COLM 2026☆17Jun 28, 2025Updated last year
- ☆293Sep 5, 2026Updated last month
- Optimize tensor program fast with Felix, a gradient descent autotuner.☆33Mar 5, 2026Updated 7 months ago
- CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation☆1,339Jul 8, 2026Updated 3 months ago
- ☆156Aug 18, 2025Updated last year
- Expert Specialization MoE Solution based on CUTLASS☆27Apr 14, 2026Updated 5 months ago
- A reference implementation of the Mind Mappings Framework.☆30Dec 2, 2021Updated 4 years ago
- Automated CUDA kernel performance diagnostics from NVIDIA Nsight Compute (NCU) CSV exports.☆35Mar 18, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- This repository provides the official implementation of QSVD, a method for efficient low-rank approximation that unifies Query-Key-Value …☆29May 16, 2026Updated 4 months ago
- [MLSys 26] 🥇 Solution for Gated Delta Net Track of MLSys 26 Flash infer competition☆36May 22, 2026Updated 4 months ago
- ☆98Feb 5, 2026Updated 8 months ago
- [MLSys 2022] "BNS-GCN: Efficient Full-Graph Training of Graph Convolutional Networks with Partition-Parallelism and Random Boundary Node …☆56Oct 6, 2023Updated 3 years ago
- Ship correct and fast LLM kernels to PyTorch☆155Jan 14, 2026Updated 8 months ago
- Sample Codes using NVSHMEM on Multi-GPU☆30Jan 22, 2023Updated 3 years ago
- Benchmarking Open-Ended Inference Optimization by AI Agents☆47Jul 6, 2026Updated 3 months ago