Shared Middle-Layer for Triton Compilation
☆34Aug 14, 2026Updated last month
Alternatives and similar repositories for triton-shared
Users that are interested in triton-shared are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆51Updated this week
- Intel® Tensor Processing Primitives extension for Pytorch*☆19Aug 24, 2026Updated 3 weeks ago
- Shared Middle-Layer for Triton Compilation☆347Dec 5, 2025Updated 9 months ago
- The docs repository of Pulsar2 which is AXera's SoC 2rd AI toolchain. Such as AX650A, AX650N☆19Sep 3, 2026Updated 2 weeks ago
- [HotStorage '24] Can ZNS SSDs be Better Storage Devices for Persistent Cache?☆13Jun 14, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ASTER 💫 : Assembly Tooling and Representations☆34Jul 1, 2026Updated 2 months ago
- Torq compiler sources☆56Jul 28, 2026Updated last month
- MLIR-based partitioning system☆210Updated this week
- Daisytuner Optimizing Compiler Collection (docc)☆22Updated this week
- ☆22Sep 27, 2022Updated 3 years ago
- FlagTree is a unified compiler supporting multiple AI chip backends for custom Deep Learning operations, which is forked from triton-lang…☆352Updated this week
- Some knowledge about riscv rvv1.0,include rvv intrinsic examples☆18Dec 30, 2025Updated 8 months ago
- A programming language with region-based memory management☆35Updated this week
- triton for dsa☆69Aug 25, 2026Updated 3 weeks ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference☆19Mar 30, 2025Updated last year
- A signal processing library in Rust, with the goal of being a decent alternative to Matlab's Signal Processing Toolbox and scipy.signal☆23Aug 7, 2026Updated last month
- low level kernels to benchmark peak compute, cache bandwidth on various levels, memory bandwidth, and some basic compute routines☆10Apr 29, 2026Updated 4 months ago
- Distributed machine learning platform☆13Aug 20, 2015Updated 11 years ago
- Hands-On Practical MLIR Tutorial☆835Oct 20, 2023Updated 2 years ago
- Multi-stage LLM agent pipeline for optimizing Triton kernels on Intel XPU — from analysis to autotuning.☆21Sep 11, 2026Updated last week
- Generator for MLIR files from known front-ends☆17Oct 31, 2023Updated 2 years ago
- A lightweight MLIR Python frontend with support for PyTorch☆28Sep 3, 2024Updated 2 years ago
- 🚧 A work-in-progress GLSL compiler targeting SPIR-V mlir 🚧☆22Oct 18, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A Python compiler design toolkit.☆590Updated this week
- TPP experimentation on MLIR for linear algebra☆165Updated this week
- An experimental CPU backend for Triton☆213Updated this week
- This repository contains companion software for the Colfax Research paper "Categorical Foundations for CuTe Layouts".☆145Aug 15, 2026Updated last month
- Enterprise Firmware platform development☆15Jun 6, 2021Updated 5 years ago
- Hardware implementation of an OmniXtend Memory Endpoint/Lowest Point of Coherence.☆20Jan 29, 2026Updated 7 months ago
- Exercises for Learning MLIR (Originally written for PPoPP 2026)☆109Jul 21, 2026Updated last month
- ☆20Jul 3, 2026Updated 2 months ago
- UCAS网络登录☆13Nov 17, 2018Updated 7 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A fast alternative to the standard C/C++ pow() function. With adjustable accuracy-space tradeoff.☆14Jul 12, 2013Updated 13 years ago
- Word2Vec 任务的并行计算实现☆11Sep 11, 2017Updated 9 years ago
- 【2024年新版】国科大 陈云霁 智能计算系统AICS实验代码☆13May 31, 2024Updated 2 years ago
- @ArchieMeng's prototype of a Python FFI of nihui/waifu2x-ncnn-vulkan achieved with SWIG☆12Jul 20, 2022Updated 4 years ago
- ☆10Aug 28, 2020Updated 6 years ago
- An Automated Performance Optimization Framework for P4-Programmable SmartNICs☆28Nov 18, 2023Updated 2 years ago
- Hexagon-MLIR is a compiler toolchain for compiling and executing AI kernels and models on Qualcomm Hexagon Neural Processing Units (NPUs)…☆231Updated this week