Agents, and RL environment, for optimizing GPU kernels on AMD ROCm using LLM agents. Benchmarks LLM serving workloads end-to-end, profiles bottleneck kernels, optimizes them via Claude Code or Codex, and scores on compilation, correctness, and speedup.
☆75Aug 8, 2026Updated this week
Alternatives and similar repositories for Apex
Users that are interested in Apex are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A lightweight, general-purpose framework for evaluating GPU kernel and benchmark.☆77Updated this week
- Jido implementation of Managed Agents on Phoenix☆23Aug 1, 2026Updated last week
- HRX: Hip Runtime Extended☆20Updated this week
- TreeFuser is a tool that perform traversals fusion for recursive tree traversals written in subset of the c++ language.☆11Aug 13, 2023Updated 2 years ago
- ☆15Jul 20, 2026Updated 3 weeks ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆19Mar 28, 2026Updated 4 months ago
- A high-performance acceleration library dedicated to large-scale model training on AMD GPUs☆68Updated this week
- Official implementation of CoNSAL for analytical Lyapunov function discovery☆12Jun 26, 2024Updated 2 years ago
- A DP beam-search extension of Mitchell Stern's span-based neural constituency parser☆11Aug 24, 2022Updated 3 years ago
- ☆30Updated this week
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆374Jul 9, 2026Updated last month
- Dashboard for InferenceX™, Open Source Continuous Inference☆37Updated this week
- Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.☆123Apr 17, 2026Updated 3 months ago
- Port of the LLVM compiler infrastructure to the time-predictable processor Patmos☆15Apr 2, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Reviving Any-Order Autoregressive Models via Principled Parallel Sampling and Speculative Decoding☆16Nov 16, 2025Updated 8 months ago
- ☆10Apr 13, 2023Updated 3 years ago
- Block Diffusion Trainer☆15Jul 10, 2025Updated last year
- Python Laboratory for Dislocation Dynamics☆13Jun 8, 2026Updated 2 months ago
- A hierarchical matrix C/C++ library☆27Updated this week
- ☆15Oct 16, 2020Updated 5 years ago
- repository for slides and code examples for my MeetingCpp talk in 2015☆12Jan 21, 2016Updated 10 years ago
- A lightweight triton-based General Matrix Multiplication (GEMM) library.☆67Jul 21, 2026Updated 2 weeks ago
- AI Tensor Engine for ROCm☆523Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- POMDP wrappers for OpenAI Gym☆15Nov 4, 2019Updated 6 years ago
- An unbounded n-gram language model on Tiny Shakespeare☆22Jan 21, 2026Updated 6 months ago
- A library containing general purpose Python utils.☆14Feb 22, 2023Updated 3 years ago
- ☆10Nov 25, 2021Updated 4 years ago
- Logger for MPI communication☆28Jul 12, 2023Updated 3 years ago
- Heterogeneous simulator for DECADES Project☆32May 23, 2024Updated 2 years ago
- An interactive web-based tool for exploring intermediate representations of PyTorch and Triton models☆49Jan 23, 2026Updated 6 months ago
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆166May 28, 2026Updated 2 months ago
- My final project, Snow Simulation, for Prof. Lingqi Yan's online open course games 101-Intro to Modern Computer Graphics☆12Mar 12, 2021Updated 5 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆15Jul 25, 2024Updated 2 years ago
- a benchmark to evaluate the situated inductive reasoning☆18Jan 7, 2025Updated last year
- A Python DSL for Apple Metal GPU compute☆25Jul 23, 2026Updated 2 weeks ago
- Heterogeneous Accelerated Computed Cluster (HACC) Resources Page☆22Jun 30, 2026Updated last month
- We introduce UltraLLaDA , a scaled variant of LLaDA-8B-Base that extends the context length up to 128K tokens with light-weight post-trai…☆15Oct 23, 2025Updated 9 months ago
- A library of fast and accurate low fidelity dynamic models for applications in robotics☆14Jul 12, 2024Updated 2 years ago
- MPI Code Generation through Domain-Specific Language Models☆16Nov 19, 2024Updated last year