A subset of PyTorch's neural network modules, written in Python using OpenAI's Triton.
☆605May 13, 2026Updated 2 months ago
Alternatives and similar repositories for attorch
Users that are interested in attorch are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Cataloging released Triton kernels.☆310Sep 9, 2025Updated 10 months ago
- Experiment of using Tangent to autodiff triton☆81Jan 22, 2024Updated 2 years ago
- Puzzles for learning Triton☆2,540Apr 1, 2026Updated 3 months ago
- Tile primitives for speedy kernels☆3,563Jul 13, 2026Updated last week
- Applied AI experiments and examples for PyTorch☆321Aug 22, 2025Updated 11 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Efficient Triton Kernels for LLM Training☆6,533Updated this week
- Triton-based implementation of Sparse Mixture of Experts.☆281Oct 3, 2025Updated 9 months ago
- Fast low-bit matmul kernels in Triton☆477Jul 15, 2026Updated last week
- A collection of memory efficient attention operators implemented in the Triton language.☆301Updated this week
- ☆115Mar 12, 2026Updated 4 months ago
- Collection of kernels written in Triton language☆200Jan 27, 2026Updated 5 months ago
- extensible collectives library in triton☆97Mar 31, 2025Updated last year
- ☆350Jul 16, 2026Updated last week
- Helpful tools and examples for working with flex-attention☆1,212Updated this week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Ring attention implementation with flash attention☆1,037Sep 10, 2025Updated 10 months ago
- Accelerated First Order Parallel Associative Scan☆198Jan 7, 2026Updated 6 months ago
- Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.☆362Updated this week
- A PyTorch native platform for training generative AI models☆5,559Updated this week
- A Quirky Assortment of CuTe Kernels☆1,070Updated this week
- Kernl lets you run PyTorch transformer models several times faster on GPU with a single line of code, and is designed to be easily hackab…☆1,585Jan 28, 2026Updated 5 months ago
- Shared Middle-Layer for Triton Compilation☆340Dec 5, 2025Updated 7 months ago
- PyTorch native quantization and sparsity for training and inference☆2,914Updated this week
- Transformers components but in Triton☆34May 9, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Mirage Persistent Kernel: Compiling LLMs into a MegaKernel☆2,390Updated this week
- ☆113Aug 26, 2024Updated last year
- UNet diffusion model in pure CUDA☆661Jun 28, 2024Updated 2 years ago
- FlexAttention w/ FlashAttention3 Support☆27Oct 5, 2024Updated last year
- Annotated version of the Mamba paper☆501Feb 27, 2024Updated 2 years ago
- 🚀 Efficient implementations for emerging model architectures☆5,414Updated this week
- A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.☆913Updated this week
- ☆124May 28, 2024Updated 2 years ago
- Development repository for the Triton language and compiler☆19,782Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Framework to reduce autotune overhead to zero for well known deployments.☆102Sep 19, 2025Updated 10 months ago
- Triton kernels for Flux☆23Jul 7, 2025Updated last year
- Implementation of a Transformer, but completely in Triton☆279Apr 5, 2022Updated 4 years ago
- GPTQ inference Triton kernel☆322May 18, 2023Updated 3 years ago
- TensorDict is a pytorch dedicated tensor container.☆1,034Updated this week
- Flash Attention in ~100 lines of CUDA (forward pass only)☆1,173Dec 30, 2024Updated last year
- ☆19Dec 4, 2025Updated 7 months ago