Shared Middle-Layer for Triton Compilation
☆36Sep 24, 2026Updated 2 weeks ago
Alternatives and similar repositories for triton-shared
Users that are interested in triton-shared are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆53Updated this week
- Intel® Tensor Processing Primitives extension for Pytorch*☆20Aug 24, 2026Updated last month
- ASTER 💫 : Assembly Tooling and Representations☆34Jul 1, 2026Updated 3 months ago
- Torq compiler sources☆57Sep 23, 2026Updated 2 weeks ago
- MLIR-based partitioning system☆213Updated this week
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Daisytuner Optimizing Compiler Collection (docc)☆23Updated this week
- FlagTree is a unified compiler supporting multiple AI chip backends for custom Deep Learning operations, which is forked from triton-lang…☆358Updated this week
- Some knowledge about riscv rvv1.0,include rvv intrinsic examples☆18Dec 30, 2025Updated 9 months ago
- A programming language with region-based memory management☆34Updated this week
- triton for dsa☆69Sep 20, 2026Updated 2 weeks ago
- low level kernels to benchmark peak compute, cache bandwidth on various levels, memory bandwidth, and some basic compute routines☆10Apr 29, 2026Updated 5 months ago
- MLIR 中文文档☆25Dec 1, 2025Updated 10 months ago
- Distributed machine learning platform☆13Aug 20, 2015Updated 11 years ago
- Hands-On Practical MLIR Tutorial☆847Oct 20, 2023Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- This repository contains the figures, tables and source code in the ICS'24 paper: "Accelerated Auto-Tuning of GPU Kernels for Tensor Comp…☆10Dec 5, 2024Updated last year
- Multi-stage LLM agent pipeline for optimizing AI kernels on Intel GPU — from analysis to autotuning.☆22Updated this week
- Generator for MLIR files from known front-ends☆17Oct 31, 2023Updated 2 years ago
- 🚧 A work-in-progress GLSL compiler targeting SPIR-V mlir 🚧☆22Oct 18, 2024Updated last year
- A Python compiler design toolkit.☆594Updated this week
- TPP experimentation on MLIR for linear algebra☆165Updated this week
- An experimental CPU backend for Triton☆214Updated this week
- Enterprise Firmware platform development☆15Jun 6, 2021Updated 5 years ago
- This repository contains companion software for the Colfax Research paper "Categorical Foundations for CuTe Layouts".☆147Aug 15, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Exercises for Learning MLIR (Originally written for PPoPP 2026)☆111Jul 21, 2026Updated 2 months ago
- UCAS网络登录☆13Nov 17, 2018Updated 7 years ago
- ☆13Nov 1, 2021Updated 4 years ago
- Word2Vec 任务的并行计算实现☆11Sep 11, 2017Updated 9 years ago
- 【2024年新版】国科大 陈云霁 智能计算系统AICS实验代码☆13May 31, 2024Updated 2 years ago
- OSDI 2023 Welder, deeplearning compiler☆36Nov 24, 2023Updated 2 years ago
- Ship correct and fast LLM kernels to PyTorch☆155Jan 14, 2026Updated 8 months ago
- Hexagon-MLIR is a compiler toolchain for compiling and executing AI kernels and models on Qualcomm Hexagon Neural Processing Units (NPUs)…☆257Updated this week
- ☆12May 18, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Wave: Python Domain-Specific Language for High Performance Machine Learning☆62Jun 29, 2026Updated 3 months ago
- The Torch-MLIR project aims to provide first class support from the PyTorch ecosystem to the MLIR ecosystem.☆1,932Updated this week
- A collection of tools and libs for C3 language☆14Sep 8, 2026Updated last month
- Unified open-source repository for OmniXtend protocol implementations in C, Verilog, and Chisel, supporting host and memory roles.☆18Apr 15, 2026Updated 5 months ago
- LLMA = LLM + Arithmetic coder, which use LLM to do insane text data compression. LLMA=大模型+算术编码,它能使用LLM对文本数据进行暴力的压缩,达到极高的压缩率。☆22Nov 24, 2024Updated last year
- MLIRX is now defunct. Please see PolyBlocks - https://docs.polymagelabs.com☆39Dec 1, 2023Updated 2 years ago
- ☆13Sep 19, 2024Updated 2 years ago