Next-Generation AI-Assisted Kernel Engineering for Multi-Chip Systems
☆72Jul 18, 2026Updated 3 weeks ago
Alternatives and similar repositories for KernelGen
Users that are interested in KernelGen are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FlagTree is a unified compiler supporting multiple AI chip backends for custom Deep Learning operations, which is forked from triton-lang…☆318Updated this week
- FlagOS skills for model deployment, HW adaptation, train&infer, eval, kernel dev and perf tuning☆19Jul 18, 2026Updated 3 weeks ago
- A vLLM plugin built on the FlagOS unified multi-chip backend.☆79Updated this week
- FlagGems is an operator library for large language models implemented in the Triton Language.☆1,073Updated this week
- Review automated kernel generation in the era of LLMs☆287Jun 25, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch…☆130Jun 29, 2026Updated last month
- ☆17Jun 10, 2025Updated last year
- MultiArchKernelBench: A Multi-Platform Benchmark for Kernel Generation☆66Jul 8, 2026Updated last month
- ☆371Updated this week
- ☆16Jul 18, 2026Updated 3 weeks ago
- KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels☆46Jun 1, 2026Updated 2 months ago
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆511Jul 15, 2026Updated 3 weeks ago
- TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators☆138Jun 14, 2025Updated last year
- ☆102Nov 22, 2025Updated 8 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Tmux sidebar for vibe coding. Manage sessions and monitor agents at a glance☆15Updated this week
- ☆28Jun 18, 2026Updated last month
- ☆181May 24, 2026Updated 2 months ago
- Code for Incorporating Relevance Feedback for Information-Seeking Retrieval using Few-Shot Document Re-Ranking, EMNLP 2022, https://aclan…☆14Mar 30, 2026Updated 4 months ago
- A CUDA kernel optimization toolkit for validation, benchmarking, Nsight Compute profiling, bottleneck analysis, and iterative tuning. It …☆195Apr 22, 2026Updated 3 months ago
- Skills for writing tilelang and debugging with CUDA toolkits.☆133May 20, 2026Updated 2 months ago
- ⌥ AI Coding agent for the terminal — hash-anchored edits, optimized tool harness, LSP, Python, browser, subagents, and agentized workflo…☆36Jun 19, 2026Updated last month
- [IJCAI 2024] CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning☆26Feb 1, 2024Updated 2 years ago
- Agentic Kernel Optimization for All — automated GPU kernel optimization for any kernel, any hardware, any language☆347May 31, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆33Mar 12, 2026Updated 5 months ago
- ☆245Updated this week
- NCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.☆24Apr 10, 2026Updated 4 months ago
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.☆1,515Mar 19, 2026Updated 4 months ago
- DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery☆23Sep 24, 2025Updated 10 months ago
- Building the Virtuous Cycle for AI-driven LLM Systems☆265May 1, 2026Updated 3 months ago
- Benchmark PyTorch Custom Operators☆14Jul 6, 2023Updated 3 years ago
- 🔥 LLM-powered GPU kernel synthesis: Train models to convert PyTorch ops into optimized Triton kernels via SFT+RL. Multi-turn compilation…☆149Nov 10, 2025Updated 9 months ago
- A benchmark of real-world DL kernel problems☆277Jul 15, 2026Updated 3 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆13Jan 7, 2025Updated last year
- RLVR training for LLM in CUDA/C++☆42Jul 25, 2026Updated 2 weeks ago
- ☆13Sep 8, 2024Updated last year
- An OpenCL-Based FPGA Accelerator for Compressed YOLOv2☆38May 27, 2021Updated 5 years ago
- [ICML 2026]☆18Jul 4, 2026Updated last month
- A Symbolic Emulator for Shuffle Synthesis on the NVIDIA PTX Code☆16Mar 19, 2023Updated 3 years ago
- The code repository for "Multi-layer Rehearsal Feature Augmentation for Class-Incremental Learning" (ICML24)☆12Jun 7, 2024Updated 2 years ago