AI Tensor Engine for ROCm
☆522Aug 9, 2026Updated this week
Alternatives and similar repositories for aiter
Users that are interested in aiter are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AiTer Optimized Model☆151Updated this week
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆260Updated this week
- Modular RDMA Interface☆165Updated this week
- [DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror☆541Updated this week
- amdgpu example code in hip/asm☆67Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming☆195Updated this week
- Fast and Furious AMD Kernels☆452Updated this week
- ☆30Updated this week
- A tool for generating information about the matrix multiplication instructions in AMD Radeon™ and AMD Instinct™ accelerators☆143Apr 10, 2026Updated 4 months ago
- ☆74Updated this week
- super repo for rocm libraries☆398Updated this week
- Ahead of Time (AOT) Triton Math Library☆100Jul 23, 2026Updated 2 weeks ago
- A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs☆119Updated this week
- Ongoing research training transformer models at scale☆43Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- super repo for rocm systems projects☆459Updated this week
- Fast and memory-efficient exact attention☆237Jul 16, 2026Updated 3 weeks ago
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆145Updated this week
- [DEPRECATED] Moved to ROCm/rocm-libraries repo☆114Jul 31, 2026Updated last week
- ☆69Updated this week
- A high-throughput and memory-efficient inference and serving engine for LLMs☆123Updated this week
- Generating Efficient AI-Centric Kernels☆145Updated this week
- MAD (Model Automation and Dashboarding)☆40Updated this week
- Development repository for the Triton language and compiler☆146Updated this week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- AMD's graph optimization engine.☆319Updated this week
- QuickReduce is a performant all-reduce library designed for AMD ROCm that supports inline compression.☆38Aug 29, 2025Updated 11 months ago
- A high-performance acceleration library dedicated to large-scale model training on AMD GPUs☆68Updated this week
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆166May 28, 2026Updated 2 months ago
- FlashInfer: Kernel Library for LLM Serving☆6,133Updated this week
- The HIP Environment and ROCm Kit - A lightweight open source build system for HIP and ROCm☆1,196Updated this week
- Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200…☆1,346Updated this week
- Distributed Compiler based on Triton for Parallel Systems☆1,512Updated this week
- [DEPRECATED] Moved to ROCm/rocm-systems repo☆92Jul 14, 2026Updated 3 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A lightweight triton-based General Matrix Multiplication (GEMM) library.☆67Jul 21, 2026Updated 2 weeks ago
- [DEPRECATED] Moved to ROCm/rocm-libraries repo☆141Jul 24, 2026Updated 2 weeks ago
- Documentation for vLLM Dev Channel releases☆10Dec 5, 2024Updated last year
- ☆125May 19, 2025Updated last year
- CUDA Templates and Python DSLs for High-Performance Linear Algebra☆10,217Updated this week
- Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels☆7,167Updated this week
- Super fast FP32 matrix multiplication on RDNA3☆92Mar 30, 2025Updated last year