Agents, and RL environment, for optimizing GPU kernels on AMD ROCm using LLM agents. Benchmarks LLM serving workloads end-to-end, profiles bottleneck kernels, optimizes them via Claude Code or Codex, and scores on compilation, correctness, and speedup.
☆76Aug 27, 2026Updated this week
Alternatives and similar repositories for Apex
Users that are interested in Apex are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A lightweight, general-purpose framework for evaluating GPU kernel and benchmark.☆80Updated this week
- LLM RL envs done right, plus some training code☆17Feb 1, 2026Updated 6 months ago
- Automated bottleneck detection and solution orchestration☆23Feb 24, 2026Updated 6 months ago
- Jido implementation of Managed Agents on Phoenix☆23Updated this week
- HRX: Hip Runtime Extended☆26Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆14Jul 20, 2026Updated last month
- A high-performance acceleration library dedicated to large-scale model training on AMD GPUs☆69Updated this week
- ☆30Aug 24, 2026Updated last week
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆377Jul 9, 2026Updated last month
- Dashboard for InferenceX™, Open Source Continuous Inference | InferenceX 仪表板☆42Updated this week
- Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.☆129Apr 17, 2026Updated 4 months ago
- Designing Full Chip from IEEE Paper using AI in SKY130☆43Jul 4, 2026Updated last month
- ☆27Aug 24, 2026Updated last week
- Extended Roofline Model - LLVM source tree with additional libraries for the analysis of the dynamic execution in the interpreter☆17Jul 5, 2017Updated 9 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs☆121Updated this week
- associative floating point addition☆21Apr 30, 2024Updated 2 years ago
- Flax (Jax) implementation of DeepSeek-R1-Distill-Qwen-1.5B with weights ported from Hugging Face.☆26Feb 20, 2025Updated last year
- ☆16Oct 21, 2025Updated 10 months ago
- A hierarchical matrix C/C++ library☆26Updated this week
- repository for slides and code examples for my MeetingCpp talk in 2015☆12Jan 21, 2016Updated 10 years ago
- A lightweight triton-based General Matrix Multiplication (GEMM) library.☆67Jul 21, 2026Updated last month
- ☆12Oct 19, 2021Updated 4 years ago
- AI Tensor Engine for ROCm☆543Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- This is master framework to automate Web, mobile, desktop and api application using Selenium, Appium, WinappDriver, RestAssured along wit…☆11Nov 29, 2023Updated 2 years ago
- Generating Efficient AI-Centric Kernels☆163Updated this week
- Winner 🏆 (Agent-only) MLSys 2026 - FlashInfer AI Kernel Generation Contest for the DeepSeek Sparse Attention (DSA) track with an average…☆157Aug 21, 2026Updated last week
- Notes and artifacts from the ONNX steering committee☆29Updated this week
- ☆10Nov 6, 2024Updated last year
- uServices - Open Vehicle Interfaces☆13Aug 2, 2024Updated 2 years ago
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆166May 28, 2026Updated 3 months ago
- ☆15Jul 25, 2024Updated 2 years ago
- A lightweight computational physics framework, based on the organization of turboWAVE. Implements a "Simulation, PhysicsModule, ComputeTo…☆12Jul 23, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A simple OpenCL path tracer with spheres and mixed glossy/specular + diffuse + emission material☆12Feb 13, 2014Updated 12 years ago
- a benchmark to evaluate the situated inductive reasoning☆19Jan 7, 2025Updated last year
- A Python DSL for Apple Metal GPU compute☆25Jul 23, 2026Updated last month
- This repository contains coding interviews that I have encountered in company interviews☆13Oct 2, 2020Updated 5 years ago
- Heterogeneous Accelerated Computed Cluster (HACC) Resources Page☆22Jun 30, 2026Updated 2 months ago
- Monorepo for osu! Stamina Trainer.☆13Updated this week
- chipStar is a tool for compiling and running HIP/CUDA on SPIR-V via OpenCL or Level Zero APIs.☆372Updated this week