Shared Middle-Layer for Triton Compilation
☆32Aug 14, 2026Updated 2 weeks ago
Alternatives and similar repositories for triton-shared
Users that are interested in triton-shared are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆48Updated this week
- Intel® Tensor Processing Primitives extension for Pytorch*☆19Updated this week
- Shared Middle-Layer for Triton Compilation☆344Dec 5, 2025Updated 8 months ago
- [HotStorage '24] Can ZNS SSDs be Better Storage Devices for Persistent Cache?☆13Jun 14, 2024Updated 2 years ago
- ASTER 💫 : Assembly Tooling and Representations☆34Jul 1, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- MLIR-based partitioning system☆207Updated this week
- ☆22Sep 27, 2022Updated 3 years ago
- Daisytuner Optimizing Compiler Collection (docc)☆22Updated this week
- FlagTree is a unified compiler supporting multiple AI chip backends for custom Deep Learning operations, which is forked from triton-lang…☆322Updated this week
- Some knowledge about riscv rvv1.0,include rvv intrinsic examples☆16Dec 30, 2025Updated 8 months ago
- triton for dsa☆69Updated this week
- InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference☆18Mar 30, 2025Updated last year
- low level kernels to benchmark peak compute, cache bandwidth on various levels, memory bandwidth, and some basic compute routines☆10Apr 29, 2026Updated 4 months ago
- MLIR 中文文档☆23Dec 1, 2025Updated 8 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Hands-On Practical MLIR Tutorial☆826Oct 20, 2023Updated 2 years ago
- This repository contains the figures, tables and source code in the ICS'24 paper: "Accelerated Auto-Tuning of GPU Kernels for Tensor Comp…☆10Dec 5, 2024Updated last year
- Multi-stage LLM agent pipeline for optimizing Triton kernels on Intel XPU — from analysis to autotuning.☆21Updated this week
- Generator for MLIR files from known front-ends☆17Oct 31, 2023Updated 2 years ago
- A lightweight MLIR Python frontend with support for PyTorch☆28Sep 3, 2024Updated last year
- 🚧 A work-in-progress GLSL compiler targeting SPIR-V mlir 🚧☆22Oct 18, 2024Updated last year
- TPP experimentation on MLIR for linear algebra☆162Updated this week
- A Python compiler design toolkit.☆582Updated this week
- An experimental CPU backend for Triton☆211Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Enterprise Firmware platform development☆15Jun 6, 2021Updated 5 years ago
- Hardware implementation of an OmniXtend Memory Endpoint/Lowest Point of Coherence.☆20Jan 29, 2026Updated 7 months ago
- Exercises for Learning MLIR (Originally written for PPoPP 2026)☆109Jul 21, 2026Updated last month
- ☆13Nov 1, 2021Updated 4 years ago
- 【2024年新版】国科大 陈云霁 智能计算系统AICS实验代码☆13May 31, 2024Updated 2 years ago
- OSDI 2023 Welder, deeplearning compiler☆35Nov 24, 2023Updated 2 years ago
- @ArchieMeng's prototype of a Python FFI of nihui/waifu2x-ncnn-vulkan achieved with SWIG☆12Jul 20, 2022Updated 4 years ago
- Ship correct and fast LLM kernels to PyTorch☆154Jan 14, 2026Updated 7 months ago
- Hexagon-MLIR is a compiler toolchain for compiling and executing AI kernels and models on Qualcomm Hexagon Neural Processing Units (NPUs)…☆205Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- An Automated Performance Optimization Framework for P4-Programmable SmartNICs☆28Nov 18, 2023Updated 2 years ago
- Wave: Python Domain-Specific Language for High Performance Machine Learning☆57Jun 29, 2026Updated 2 months ago
- ☆12May 18, 2024Updated 2 years ago
- dump and replace shaders of any OpenGL or Vulkan application☆31May 24, 2018Updated 8 years ago
- The Torch-MLIR project aims to provide first class support from the PyTorch ecosystem to the MLIR ecosystem.☆1,894Updated this week
- LLMA = LLM + Arithmetic coder, which use LLM to do insane text data compression. LLMA=大模型+算术编码,它能使用LLM对文本数据进行暴力的压缩,达到极高的压缩率。☆22Nov 24, 2024Updated last year
- PyDTNN - Python Distributed Training of Neural Networks☆14Updated this week