An MLIR-based compiler that takes GPU kernels and compiles them to real hardware instructions. Interactive web visualizer included.
☆144Mar 21, 2026Updated 5 months ago
Alternatives and similar repositories for tiny-gpu-compiler
Users that are interested in tiny-gpu-compiler are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Tutorial on building a gpu compiler backend in LLVM☆67Jan 11, 2025Updated last year
- python driver and runtime for tenstorrent blackhole cards☆19Aug 12, 2026Updated 2 weeks ago
- ☆17Mar 29, 2026Updated 5 months ago
- A visualizer for the ROCm Profiler Tools☆26Updated this week
- A double JIT VM☆23Jul 9, 2026Updated last month
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- QuickReduce is a performant all-reduce library designed for AMD ROCm that supports inline compression.☆38Aug 29, 2025Updated last year
- A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.☆81Feb 18, 2026Updated 6 months ago
- LLVM Code Generation, published by Packt☆284May 14, 2026Updated 3 months ago
- An online tutorial to make MLIR more beginner friendly with an end-to-end deep learning compiler pipeline☆62Jun 8, 2026Updated 2 months ago
- FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kerne…☆269Updated this week
- SBLP 2025 MLIR Tutorial☆77Mar 25, 2026Updated 5 months ago
- A minimal tensor processing unit (TPU), inspired by Google's TPU V2 and V1☆1,369Apr 3, 2026Updated 4 months ago
- Reverse engineering NVIDIA SASS instruction dictionary, kernel audits and pattern recognition across GPU architectures.☆335May 18, 2026Updated 3 months ago
- a whirlwind tour to deep learning and deep learning systems☆82Updated this week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Monte Carlo neutron transport in C99. GPU via Booth (AMD MI300X, NVIDIA RTX). ENDF/B-VII.1 nuclear data. Validated against ICSBEP benchma…☆15Aug 21, 2026Updated last week
- Open-source CUDA, Triton and HIP compiler targeting multiple GPU and CPU architectures.☆1,732Aug 24, 2026Updated last week
- Exercises for Learning MLIR (Originally written for PPoPP 2026)☆109Jul 21, 2026Updated last month
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆377Jul 9, 2026Updated last month
- Build an LLM from scratch with MAX☆75Updated this week
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆240Jun 29, 2026Updated 2 months ago
- A curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.☆2,190Aug 23, 2026Updated last week
- A tool for generating information about the matrix multiplication instructions in AMD Radeon™ and AMD Instinct™ accelerators☆143Apr 10, 2026Updated 4 months ago
- Hexagon-MLIR is a compiler toolchain for compiling and executing AI kernels and models on Qualcomm Hexagon Neural Processing Units (NPUs)…☆205Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- TensaLang is a Tensor-first programming language, compiler, and runtime that let you write the Model’s inference engine (e.g. LLMs) and s…☆76Feb 20, 2026Updated 6 months ago
- ☆125Aug 8, 2026Updated 3 weeks ago
- A performance portable numerical relativity code built on AMReX (successor to GRChombo)☆18Updated this week
- GPU kernel benchmarking☆47Jun 10, 2026Updated 2 months ago
- ☆46May 24, 2025Updated last year
- Fast and Furious AMD Kernels☆461Updated this week
- Alpha64 R10000 Two-Way Superscalar Processor☆12May 6, 2019Updated 7 years ago
- ☆27Jun 5, 2025Updated last year
- Linux kernel Type-1 hypervisor with modular VMX, SVM, ARM EL2, and RISC-V H support☆30Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A toy compiler for NumPy array expressions that uses e-graphs and MLIR☆127Aug 24, 2026Updated last week
- ☆31Mar 22, 2026Updated 5 months ago
- Interactive version of the CuTe layout paper☆57Apr 14, 2026Updated 4 months ago
- High-Performance FP32 GEMM on CUDA devices☆126Jan 21, 2025Updated last year
- High-performance GPU kernels for LLM inference in OpenAI Triton. Fused RMSNorm, SwiGLU, INT8 GEMM with benchmarks and roofline analysis.☆42Jul 22, 2026Updated last month
- A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.☆930Updated this week
- A lightweight triton-based General Matrix Multiplication (GEMM) library.☆67Jul 21, 2026Updated last month