An MLIR-based compiler that takes GPU kernels and compiles them to real hardware instructions. Interactive web visualizer included.
☆139Mar 21, 2026Updated 4 months ago
Alternatives and similar repositories for tiny-gpu-compiler
Users that are interested in tiny-gpu-compiler are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Tutorial on building a gpu compiler backend in LLVM☆62Jan 11, 2025Updated last year
- python driver and runtime for tenstorrent blackhole cards☆16Updated this week
- KV Cache & LoRA for minGPT☆61Mar 4, 2026Updated 4 months ago
- minimal compiler☆24Feb 19, 2026Updated 5 months ago
- A visualizer for the ROCm Profiler Tools☆24Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A double JIT VM☆22Jul 9, 2026Updated last week
- World's first Nintendo 3DS emulator for Apple devices based on Citra.☆18Apr 7, 2023Updated 3 years ago
- QuickReduce is a performant all-reduce library designed for AMD ROCm that supports inline compression.☆38Aug 29, 2025Updated 10 months ago
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆79Feb 18, 2026Updated 5 months ago
- LLVM Code Generation, published by Packt☆272May 14, 2026Updated 2 months ago
- An online tutorial to make MLIR more beginner friendly with an end-to-end deep learning compiler pipeline☆55Jun 8, 2026Updated last month
- FlyDSL is the Python front‑end of the project: Flexible LaYout DSL.☆237Updated this week
- SBLP 2025 MLIR Tutorial☆75Mar 25, 2026Updated 3 months ago
- A minimal tensor processing unit (TPU), inspired by Google's TPU V2 and V1☆1,349Apr 3, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Reverse engineering NVIDIA SASS instruction dictionary, kernel audits and pattern recognition across GPU architectures.☆311May 18, 2026Updated 2 months ago
- a whirlwind tour to deep learning and deep learning systems☆81Updated this week
- Monte Carlo neutron transport in C99. GPU via Booth (AMD MI300X, NVIDIA RTX). ENDF/B-VII.1 nuclear data. Validated against ICSBEP benchma…☆15Jun 5, 2026Updated last month
- Open-source CUDA, Triton and HIP compiler targeting multiple GPU and CPU architectures.☆1,717Updated this week
- hands on model tuning with TVM and profile it on a Mac M1, x86 CPU, and GTX-1080 GPU.☆51Jun 15, 2023Updated 3 years ago
- This repository contains my work from the VLSI System Design (VSD) Workshop on 7nm FinFET Circuit Design and Characterization using the A…☆16Sep 9, 2025Updated 10 months ago
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆367Jul 9, 2026Updated last week
- Exercises for Learning MLIR (Originally written for PPoPP 2026)☆106Feb 5, 2026Updated 5 months ago
- Build an LLM from scratch with MAX☆64Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆231Jun 29, 2026Updated 3 weeks ago
- 🍒 Cherry programming language☆17Sep 18, 2024Updated last year
- Tensor library & inference framework for machine learning☆118Oct 3, 2025Updated 9 months ago
- A curriculum for learning about gpu performance engineering, from scratch to what the frontier AI labs do☆1,256Apr 27, 2026Updated 2 months ago
- Hexagon-MLIR is a compiler toolchain for compiling and executing AI kernels and models on Qualcomm Hexagon Neural Processing Units (NPUs)…☆177Jul 2, 2026Updated 2 weeks ago
- TensaLang is a Tensor-first programming language, compiler, and runtime that let you write the Model’s inference engine (e.g. LLMs) and s…☆77Feb 20, 2026Updated 5 months ago
- GPU kernel benchmarking☆47Jun 10, 2026Updated last month
- ☆46May 24, 2025Updated last year
- Fast and Furious AMD Kernels☆444Jul 10, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Native Linux KVM tool☆15May 18, 2026Updated 2 months ago
- Alpha64 R10000 Two-Way Superscalar Processor☆12May 6, 2019Updated 7 years ago
- ☆27Jun 5, 2025Updated last year
- Linux kernel Type-1 hypervisor with modular VMX, SVM, ARM EL2, and RISC-V H support☆27Updated this week
- A toy compiler for NumPy array expressions that uses e-graphs and MLIR☆122Jul 13, 2026Updated last week
- ☆30Mar 22, 2026Updated 3 months ago
- Interactive version of the CuTe layout paper☆57Apr 14, 2026Updated 3 months ago