π₯ LLM-powered GPU kernel synthesis: Train models to convert PyTorch ops into optimized Triton kernels via SFT+RL. Multi-turn compilation feedback, cross-platform NVIDIA/AMD, Kernelbook + KernelBench
β152Nov 10, 2025Updated 10 months ago
Alternatives and similar repositories for TritonForge
Users that are interested in TritonForge are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FA4-based Relative Attention Kernel developed by TML and Colfaxβ18Sep 11, 2026Updated last week
- A lightweight post-training framework for LLMs and VLMs. 51 algorithms, 38 verified models. Scales with DeepSpeed, vLLM, and Ray.β20Aug 4, 2026Updated last month
- β34Mar 12, 2026Updated 6 months ago
- β37Aug 7, 2025Updated last year
- APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation. A system-level optimization for scalable LLM traβ¦β64Oct 11, 2025Updated 11 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Allow torch tensor memory to be released and resumed laterβ275Sep 12, 2026Updated last week
- [MLSys 26] π₯ Solution for Gated Delta Net Track of MLSys 26 Flash infer competitionβ36May 22, 2026Updated 3 months ago
- Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.β2,952Updated this week
- Tinkering RLβ29Updated this week
- Synchronizing Claude Code conversations across machinesβ17Aug 31, 2026Updated 3 weeks ago
- A collection of specialized agent skills for AI infrastructure development, enabling Claude Code to write, optimize, and debug high-perfoβ¦β147Jul 9, 2026Updated 2 months ago
- Accelerating MoE with IO and Tile-aware Optimizationsβ768Aug 29, 2026Updated 3 weeks ago
- Autonomous GPU Kernel Generation & Optimization via Deep Agentsβ556Sep 8, 2026Updated last week
- [NeurIPS 2025] Scaling Speculative Decoding with Lookahead Reasoningβ69Oct 31, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Bridge Megatron-Core to Hugging Face/Reinforcement Learningβ232Jun 15, 2026Updated 3 months ago
- Train speculative decoding models effortlessly and port them smoothly to SGLang serving.β1,178Updated this week
- TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operatorsβ140Jun 14, 2025Updated last year
- FlashInfer Bench @ MLSys 2026: Building AI agents to write high performance GPU kernelsβ183Apr 26, 2026Updated 4 months ago
- β255Nov 19, 2025Updated 10 months ago
- β39Dec 14, 2025Updated 9 months ago
- Implementation for FP8/INT8 Rollout for RL training without performence drop.β308Nov 7, 2025Updated 10 months ago
- A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Trainingβ945Updated this week
- Official repository Flash Local Linear Attentionβ40May 28, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Autonomous GPU kernel optimization system driven by AI agents.β31Mar 29, 2026Updated 5 months ago
- β54May 19, 2025Updated last year
- [KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICMLβ¦β210Mar 29, 2026Updated 5 months ago
- β458Aug 26, 2026Updated 3 weeks ago
- slime is an LLM post-training framework for RL Scaling.β8,509Updated this week
- My love.β28Apr 1, 2026Updated 5 months ago
- [ASPLOS'26] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafterβ179Feb 27, 2026Updated 6 months ago
- β73Mar 24, 2026Updated 5 months ago
- KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)β1,252Mar 24, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Compact and Agent-Native MoE Training Systemβ352Updated this week
- (best/better) practices of megatron on veRL and tuning guideβ138May 12, 2026Updated 4 months ago
- FLA but cuTileβ27Apr 17, 2026Updated 5 months ago
- β861Updated this week
- A Quirky Assortment of CuTe Kernelsβ1,149Updated this week
- Kernel Design Agents (KDA) is a agent-centric workflow to write high-performance CUDA Kernels.β1,053Sep 14, 2026Updated last week
- β97Sep 15, 2025Updated last year