Official Repo of CudaForge
☆89Dec 2, 2025Updated 9 months ago
Alternatives and similar repositories for CudaForge
Users that are interested in CudaForge are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆25Feb 14, 2026Updated 7 months ago
- GraphDancer: Training LLMs to Explore and Reason over Graphs via Curriculum Reinforcement Learning☆21May 25, 2026Updated 3 months ago
- LLM4Kernel: A Survey of Large Language Models for GPU Kernel Development☆82Mar 31, 2026Updated 5 months ago
- ☆103Nov 22, 2025Updated 9 months ago
- Skills for writing tilelang and debugging with CUDA toolkits.☆144Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆556Sep 8, 2026Updated last week
- CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning☆323Nov 3, 2025Updated 10 months ago
- A benchmark of real-world DL kernel problems☆296Jul 15, 2026Updated 2 months ago
- ☆12Aug 4, 2022Updated 4 years ago
- ☆168Aug 8, 2026Updated last month
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)☆1,250Mar 24, 2026Updated 5 months ago
- Building the Virtuous Cycle for AI-driven LLM Systems☆281May 1, 2026Updated 4 months ago
- ☆16Jan 24, 2024Updated 2 years ago
- Some funny cute/cuteDSL code snippets☆35Mar 2, 2026Updated 6 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆232Aug 26, 2026Updated 3 weeks ago
- Speed of Light Analysis for ML Model Runtime☆123Sep 13, 2026Updated last week
- [KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML…☆210Mar 29, 2026Updated 5 months ago
- This is code for How Do Social Bots Participate in Misinformation Spread? A Comprehensive Dataset and Analysis☆18Nov 5, 2025Updated 10 months ago
- ☆457Aug 26, 2026Updated 3 weeks ago
- PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.☆228Dec 24, 2025Updated 8 months ago
- ☆19May 9, 2025Updated last year
- Automated High-Performance GPU Kernel Generation☆136Jun 1, 2026Updated 3 months ago
- Evaluating Large Language Models for CUDA Code Generation ComputeEval is a framework designed to generate and evaluate CUDA code from Lar…☆149Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆101Mar 31, 2026Updated 5 months ago
- ScienceMeter: Tracking Scientific Knowledge Updates in Language Models, COLM 2026☆17Jun 28, 2025Updated last year
- ☆284Sep 5, 2026Updated 2 weeks ago
- Optimize tensor program fast with Felix, a gradient descent autotuner.☆33Mar 5, 2026Updated 6 months ago
- CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation☆1,320Jul 8, 2026Updated 2 months ago
- ☆155Aug 18, 2025Updated last year
- Expert Specialization MoE Solution based on CUTLASS☆27Apr 14, 2026Updated 5 months ago
- A reference implementation of the Mind Mappings Framework.☆30Dec 2, 2021Updated 4 years ago
- Automated CUDA kernel performance diagnostics from NVIDIA Nsight Compute (NCU) CSV exports.☆34Mar 18, 2026Updated 6 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- This repository provides the official implementation of QSVD, a method for efficient low-rank approximation that unifies Query-Key-Value …☆28May 16, 2026Updated 4 months ago
- [MLSys 26] 🥇 Solution for Gated Delta Net Track of MLSys 26 Flash infer competition☆36May 22, 2026Updated 3 months ago
- ☆93Feb 5, 2026Updated 7 months ago
- [MLSys 2022] "BNS-GCN: Efficient Full-Graph Training of Graph Convolutional Networks with Partition-Parallelism and Random Boundary Node …☆56Oct 6, 2023Updated 2 years ago
- Ship correct and fast LLM kernels to PyTorch☆154Jan 14, 2026Updated 8 months ago
- Sample Codes using NVSHMEM on Multi-GPU☆30Jan 22, 2023Updated 3 years ago
- Benchmarking Open-Ended Inference Optimization by AI Agents☆44Jul 6, 2026Updated 2 months ago