Automated bottleneck detection and solution orchestration
☆23Feb 24, 2026Updated 4 months ago
Alternatives and similar repositories for intelliperf
Users that are interested in intelliperf are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- IntelliKit is a collection of intelligent tools designed to make GPU kernel development, profiling, and validation accessible to LLMs and…☆27Updated this week
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆193Updated this week
- Scale-out system monitoring☆25Updated this week
- A lightweight, general-purpose framework for evaluating GPU kernel and benchmark.☆56Updated this week
- Tensor library for machine learning☆34Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A lightweight triton-based General Matrix Multiplication (GEMM) library.☆65Updated this week
- HRX: Hip Runtime Extended☆18Updated this week
- Automating analysis from trace files☆81Updated this week
- ☆101Nov 22, 2025Updated 8 months ago
- Generating Efficient AI-Centric Kernels☆123Updated this week
- ☆62Jul 16, 2026Updated last week
- ☆30Jun 16, 2026Updated last month
- ☆10May 15, 2024Updated 2 years ago
- Header-only library of GPU-accelerated, concurrent data structures.☆12Jul 10, 2026Updated last week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆14Oct 5, 2022Updated 3 years ago
- Framework to reduce autotune overhead to zero for well known deployments.☆101Sep 19, 2025Updated 10 months ago
- Examples illustrating usage of the rocBLAS library☆17Aug 12, 2024Updated last year
- ☆16Updated this week
- ☆19Mar 29, 2026Updated 3 months ago
- Public benchmark results from Kernel Arena, a leaderboard for LLM-generated AI accelerator kernels.☆20Mar 11, 2026Updated 4 months ago
- ☆21Mar 17, 2026Updated 4 months ago
- Repository with examples and exercises for OLCF and AMD's HIP training series☆17Oct 16, 2023Updated 2 years ago
- Agents, and RL environment, for optimizing GPU kernels on AMD ROCm using LLM agents. Benchmarks LLM serving workloads end-to-end, profile…☆71Jul 16, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆30May 28, 2026Updated last month
- ☆19Nov 11, 2025Updated 8 months ago
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆146Updated this week
- AI Tensor Engine for ROCm☆497Updated this week
- ☆20May 30, 2026Updated last month
- ☆19May 9, 2025Updated last year
- ☆19Jun 6, 2025Updated last year
- NVFP4 Flash-Attention 4 on BlackWell☆30Updated this week
- Evaluating Large Language Models for CUDA Code Generation ComputeEval is a framework designed to generate and evaluate CUDA code from Lar…☆143May 19, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Parallel Programming with MPI☆25May 11, 2026Updated 2 months ago
- A Simple Network Ping and Traceroute Tool☆31Sep 24, 2015Updated 10 years ago
- ☆350Updated this week
- Distributed machine learning platform☆13Aug 20, 2015Updated 10 years ago
- LLVM/MLIR based compiler instrumentation of AMD GPU kernels☆21Jul 13, 2025Updated last year
- FLA but cuTile☆27Apr 17, 2026Updated 3 months ago
- Ring network model test to demonstrate the use of CoreNEURON☆11Jul 5, 2026Updated 2 weeks ago