Next-Generation AI-Assisted Kernel Engineering for Multi-Chip Systems
☆68Jul 18, 2026Updated this week
Alternatives and similar repositories for KernelGen
Users that are interested in KernelGen are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FlagOS skills for model deployment, HW adaptation, train&infer, eval, kernel dev and perf tuning☆19Updated this week
- FlagTree is a unified compiler supporting multiple AI chip backends for custom Deep Learning operations, which is forked from triton-lang…☆301Updated this week
- A vLLM plugin built on the FlagOS unified multi-chip backend.☆66Updated this week
- FlagGems is an operator library for large language models implemented in the Triton Language.☆1,053Updated this week
- Review automated kernel generation in the era of LLMs☆274Jun 25, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch…☆119Jun 29, 2026Updated 3 weeks ago
- ☆17Jun 10, 2025Updated last year
- MultiArchKernelBench: A Multi-Platform Benchmark for Kernel Generation☆63Jul 8, 2026Updated 2 weeks ago
- ☆312Jun 9, 2026Updated last month
- ☆16Updated this week
- Autonomous GPU Kernel Generation & Optimization via Deep Agents☆489Jul 15, 2026Updated last week
- TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators☆137Jun 14, 2025Updated last year
- ☆101Nov 22, 2025Updated 8 months ago
- Tmux sidebar for vibe coding. Manage sessions and monitor agents at a glance☆15Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆27Jun 18, 2026Updated last month
- ☆157May 24, 2026Updated last month
- Code for Incorporating Relevance Feedback for Information-Seeking Retrieval using Few-Shot Document Re-Ranking, EMNLP 2022, https://aclan…☆14Mar 30, 2026Updated 3 months ago
- A CUDA kernel optimization toolkit for validation, benchmarking, Nsight Compute profiling, bottleneck analysis, and iterative tuning. It …☆191Apr 22, 2026Updated 3 months ago
- PTO instruction set architecture☆64Updated this week
- Skills for writing tilelang and debugging with CUDA toolkits.☆133May 20, 2026Updated 2 months ago
- ⌥ AI Coding agent for the terminal — hash-anchored edits, optimized tool harness, LSP, Python, browser, subagents, and agentized workflo…☆36Jun 19, 2026Updated last month
- [IJCAI 2024] CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning☆26Feb 1, 2024Updated 2 years ago
- Agentic Kernel Optimization for All — automated GPU kernel optimization for any kernel, any hardware, any language☆326May 31, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆33Mar 12, 2026Updated 4 months ago
- ☆233Updated this week
- NCU-driven iterative optimization workflow for CUDA/CUTLASS/Triton/CuTe DSL kernels.☆23Apr 10, 2026Updated 3 months ago
- Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.☆1,475Mar 19, 2026Updated 4 months ago
- DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery☆23Sep 24, 2025Updated 9 months ago
- ChipScope / ILA using XVC (XIlinx Virtual Cable Over PCIe) with a PR (Partial Reconfiguration) design Example.☆15Jun 1, 2017Updated 9 years ago
- Building the Virtuous Cycle for AI-driven LLM Systems☆261May 1, 2026Updated 2 months ago
- Benchmark PyTorch Custom Operators☆14Jul 6, 2023Updated 3 years ago
- 🔥 LLM-powered GPU kernel synthesis: Train models to convert PyTorch ops into optimized Triton kernels via SFT+RL. Multi-turn compilation…☆146Nov 10, 2025Updated 8 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- HLS implementation of cuckoo hashing. Refer to paper : https://ieeexplore.ieee.org/document/7577355/☆14Dec 4, 2018Updated 7 years ago
- ☆13Jan 7, 2025Updated last year
- A crowd-powered database system, with SQL-like query interface, multi-goal optimization☆11Sep 4, 2017Updated 8 years ago
- RLVR training for LLM in CUDA/C++☆39Jun 8, 2026Updated last month
- An OpenCL-Based FPGA Accelerator for Compressed YOLOv2☆38May 27, 2021Updated 5 years ago
- A Symbolic Emulator for Shuffle Synthesis on the NVIDIA PTX Code☆16Mar 19, 2023Updated 3 years ago
- The code repository for "Multi-layer Rehearsal Feature Augmentation for Class-Incremental Learning" (ICML24)☆12Jun 7, 2024Updated 2 years ago