Agents, and RL environment, for optimizing GPU kernels on AMD ROCm using LLM agents. Benchmarks LLM serving workloads end-to-end, profiles bottleneck kernels, optimizes them via Claude Code or Codex, and scores on compilation, correctness, and speedup.
☆71Jul 16, 2026Updated this week
Alternatives and similar repositories for Apex
Users that are interested in Apex are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A lightweight, general-purpose framework for evaluating GPU kernel and benchmark.☆56Updated this week
- LLM RL envs done right, plus some training code☆17Feb 1, 2026Updated 5 months ago
- Automated bottleneck detection and solution orchestration☆23Feb 24, 2026Updated 4 months ago
- Jido implementation of Managed Agents on Phoenix☆23Updated this week
- HRX: Hip Runtime Extended☆18Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- TreeFuser is a tool that perform traversals fusion for recursive tree traversals written in subset of the c++ language.☆11Aug 13, 2023Updated 2 years ago
- A high-performance acceleration library dedicated to large-scale model training on AMD GPUs☆67Updated this week
- Repository for getting started with the OfficeQA Benchmark.☆162Jun 13, 2026Updated last month
- The Advocacy Platform is a cloud solution to automate the acquisition of case decisions, court hearings and location information from the…☆24Oct 9, 2023Updated 2 years ago
- ☆30Jun 16, 2026Updated last month
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆367Jul 9, 2026Updated last week
- Dashboard for InferenceX™, Open Source Continuous Inference☆36Updated this week
- Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.☆121Apr 17, 2026Updated 3 months ago
- Designing Full Chip from IEEE Paper using AI in SKY130☆23Jul 4, 2026Updated 2 weeks ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Extended Roofline Model - LLVM source tree with additional libraries for the analysis of the dynamic execution in the interpreter☆17Jul 5, 2017Updated 9 years ago
- Flax (Jax) implementation of DeepSeek-R1-Distill-Qwen-1.5B with weights ported from Hugging Face.☆26Feb 20, 2025Updated last year
- A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs☆107Updated this week
- Port of the LLVM compiler infrastructure to the time-predictable processor Patmos☆15Apr 2, 2025Updated last year
- A hierarchical matrix C/C++ library☆27Jul 10, 2026Updated last week
- A lightweight triton-based General Matrix Multiplication (GEMM) library.☆65Jun 13, 2026Updated last month
- AI Tensor Engine for ROCm☆497Updated this week
- This is master framework to automate Web, mobile, desktop and api application using Selenium, Appium, WinappDriver, RestAssured along wit…☆11Nov 29, 2023Updated 2 years ago
- Generating Efficient AI-Centric Kernels☆121Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆10Nov 25, 2021Updated 4 years ago
- High-performance GEMM kernel examples with FlyDSL on AMD GPUs.☆27Updated this week
- Logger for MPI communication☆28Jul 12, 2023Updated 3 years ago
- Heterogeneous simulator for DECADES Project☆32May 23, 2024Updated 2 years ago
- uServices - Open Vehicle Interfaces☆13Aug 2, 2024Updated last year
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆165May 28, 2026Updated last month
- ☆15Jul 25, 2024Updated last year
- Five Bottlenecks, Five Fixes: How We Avoid Leaving Training Performance on the Table☆19Mar 13, 2026Updated 4 months ago
- Python CFFI Binding around SuiteSparse:GraphBLAS☆24Apr 27, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A simple OpenCL path tracer with spheres and mixed glossy/specular + diffuse + emission material☆12Feb 13, 2014Updated 12 years ago
- a benchmark to evaluate the situated inductive reasoning☆16Jan 7, 2025Updated last year
- This repository contains coding interviews that I have encountered in company interviews☆13Oct 2, 2020Updated 5 years ago
- Heterogeneous Accelerated Computed Cluster (HACC) Resources Page☆22Jun 30, 2026Updated 3 weeks ago
- Distributed AtomSpace Network client☆19Jan 19, 2026Updated 6 months ago
- chipStar is a tool for compiling and running HIP/CUDA on SPIR-V via OpenCL or Level Zero APIs.☆364Updated this week
- ☆13Apr 16, 2025Updated last year