An MLIR-based compiler that takes GPU kernels and compiles them to real hardware instructions. Interactive web visualizer included.
☆147Mar 21, 2026Updated 6 months ago
Alternatives and similar repositories for tiny-gpu-compiler
Users that are interested in tiny-gpu-compiler are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Tutorial on building a gpu compiler backend in LLVM☆66Jan 11, 2025Updated last year
- python driver and runtime for tenstorrent blackhole cards☆20Sep 16, 2026Updated 3 weeks ago
- KV Cache & LoRA for minGPT☆61Mar 4, 2026Updated 7 months ago
- ☆17Mar 29, 2026Updated 6 months ago
- A visualizer for the ROCm Profiler Tools☆32Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- World's first Nintendo 3DS emulator for Apple devices based on Citra.☆18Apr 7, 2023Updated 3 years ago
- A double JIT VM☆23Jul 9, 2026Updated 3 months ago
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆83Feb 18, 2026Updated 7 months ago
- LLVM Code Generation, published by Packt☆287May 14, 2026Updated 4 months ago
- An online tutorial to make MLIR more beginner friendly with an end-to-end deep learning compiler pipeline☆68Jun 8, 2026Updated 4 months ago
- SBLP 2025 MLIR Tutorial☆79Mar 25, 2026Updated 6 months ago
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆290Updated this week
- A minimal tensor processing unit (TPU), inspired by Google's TPU V2 and V1☆1,393Apr 3, 2026Updated 6 months ago
- Reverse engineering NVIDIA SASS instruction dictionary, kernel audits and pattern recognition across GPU architectures.☆342May 18, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Monte Carlo neutron transport in C99. GPU via Booth (AMD MI300X, NVIDIA RTX). ENDF/B-VII.1 nuclear data. Validated against ICSBEP benchma…☆14Aug 21, 2026Updated last month
- a whirlwind tour to deep learning and deep learning systems☆95Updated this week
- Open-source CUDA, Triton and HIP compiler targeting multiple GPU and CPU architectures.☆1,752Sep 14, 2026Updated 3 weeks ago
- hands on model tuning with TVM and profile it on a Mac M1, x86 CPU, and GTX-1080 GPU.☆51Jun 15, 2023Updated 3 years ago
- Exercises for Learning MLIR (Originally written for PPoPP 2026)☆111Jul 21, 2026Updated 2 months ago
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆384Jul 9, 2026Updated 3 months ago
- Hexagon-MLIR is a compiler toolchain for compiling and executing AI kernels and models on Qualcomm Hexagon Neural Processing Units (NPUs)…☆257Updated this week
- This repository contains my work from the VLSI System Design (VSD) Workshop on 7nm FinFET Circuit Design and Characterization using the A…☆16Sep 9, 2025Updated last year
- Build an LLM from scratch with MAX☆78Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆253Jun 29, 2026Updated 3 months ago
- Tensor library & inference framework for machine learning☆119Oct 3, 2025Updated last year
- A tool for generating information about the matrix multiplication instructions in AMD Radeon™ and AMD Instinct™ accelerators☆145Sep 24, 2026Updated 2 weeks ago
- A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.☆4,099Sep 12, 2026Updated 3 weeks ago
- TensaLang is a Tensor-first programming language, compiler, and runtime that let you write the Model’s inference engine (e.g. LLMs) and s…☆77Feb 20, 2026Updated 7 months ago
- ☆148Sep 15, 2026Updated 3 weeks ago
- A performance portable numerical relativity code built on AMReX (successor to GRChombo)☆18Oct 1, 2026Updated last week
- GPU kernel benchmarking☆49Jun 10, 2026Updated 4 months ago
- Fast and Furious AMD Kernels☆473Oct 3, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Alpha64 R10000 Two-Way Superscalar Processor☆12May 6, 2019Updated 7 years ago
- ☆27Jun 5, 2025Updated last year
- A toy compiler for NumPy array expressions that uses e-graphs and MLIR☆130Updated this week
- ☆31Mar 22, 2026Updated 6 months ago
- Interactive version of the CuTe layout paper☆57Apr 14, 2026Updated 5 months ago
- High-Performance FP32 GEMM on CUDA devices☆130Jan 21, 2025Updated last year
- 🍒 Cherry programming language☆18Sep 18, 2024Updated 2 years ago