An MLIR-based compiler that takes GPU kernels and compiles them to real hardware instructions. Interactive web visualizer included.
☆145Mar 21, 2026Updated 5 months ago
Alternatives and similar repositories for tiny-gpu-compiler
Users that are interested in tiny-gpu-compiler are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Tutorial on building a gpu compiler backend in LLVM☆67Jan 11, 2025Updated last year
- minimal compiler☆24Feb 19, 2026Updated 7 months ago
- ☆17Mar 29, 2026Updated 5 months ago
- A visualizer for the ROCm Profiler Tools☆30Updated this week
- A double JIT VM☆23Jul 9, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- QuickReduce is a performant all-reduce library designed for AMD ROCm that supports inline compression.☆38Aug 29, 2025Updated last year
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆82Feb 18, 2026Updated 7 months ago
- LLVM Code Generation, published by Packt☆287May 14, 2026Updated 4 months ago
- SBLP 2025 MLIR Tutorial☆77Mar 25, 2026Updated 5 months ago
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆282Updated this week
- A minimal tensor processing unit (TPU), inspired by Google's TPU V2 and V1☆1,377Apr 3, 2026Updated 5 months ago
- Reverse engineering NVIDIA SASS instruction dictionary, kernel audits and pattern recognition across GPU architectures.☆340May 18, 2026Updated 4 months ago
- a whirlwind tour to deep learning and deep learning systems☆83Updated this week
- Monte Carlo neutron transport in C99. GPU via Booth (AMD MI300X, NVIDIA RTX). ENDF/B-VII.1 nuclear data. Validated against ICSBEP benchma…☆15Aug 21, 2026Updated 3 weeks ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Open-source CUDA, Triton and HIP compiler targeting multiple GPU and CPU architectures.☆1,749Updated this week
- hands on model tuning with TVM and profile it on a Mac M1, x86 CPU, and GTX-1080 GPU.☆51Jun 15, 2023Updated 3 years ago
- Exercises for Learning MLIR (Originally written for PPoPP 2026)☆109Jul 21, 2026Updated last month
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆382Jul 9, 2026Updated 2 months ago
- Hexagon-MLIR is a compiler toolchain for compiling and executing AI kernels and models on Qualcomm Hexagon Neural Processing Units (NPUs)…☆231Updated this week
- Build an LLM from scratch with MAX☆75Updated this week
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆243Jun 29, 2026Updated 2 months ago
- Tensor library & inference framework for machine learning☆118Oct 3, 2025Updated 11 months ago
- A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.☆3,210Sep 12, 2026Updated last week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A tool for generating information about the matrix multiplication instructions in AMD Radeon™ and AMD Instinct™ accelerators☆144Apr 10, 2026Updated 5 months ago
- ☆139Updated this week
- A performance portable numerical relativity code built on AMReX (successor to GRChombo)☆18Updated this week
- GPU kernel benchmarking☆48Jun 10, 2026Updated 3 months ago
- ☆47May 24, 2025Updated last year
- Fast and Furious AMD Kernels☆470Updated this week
- Native Linux KVM tool☆16Aug 13, 2026Updated last month
- Alpha64 R10000 Two-Way Superscalar Processor☆12May 6, 2019Updated 7 years ago
- ☆27Jun 5, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Linux kernel Type-1 hypervisor with modular VMX, SVM, ARM EL2, and RISC-V H support☆30Sep 8, 2026Updated last week
- A toy compiler for NumPy array expressions that uses e-graphs and MLIR☆129Updated this week
- ☆31Mar 22, 2026Updated 5 months ago
- Interactive version of the CuTe layout paper☆57Apr 14, 2026Updated 5 months ago
- High-Performance FP32 GEMM on CUDA devices☆128Jan 21, 2025Updated last year
- 🍒 Cherry programming language☆18Sep 18, 2024Updated 2 years ago
- High-performance GPU kernels for LLM inference in OpenAI Triton. Fused RMSNorm, SwiGLU, INT8 GEMM with benchmarks and roofline analysis.☆43Jul 22, 2026Updated last month