Triton based sparse quantization attention kernel collection
☆43Aug 29, 2025Updated last year
Alternatives and similar repositories for attention-gym
Users that are interested in attention-gym are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Distributed parallel 3D-Causal-VAE for efficient training and inference☆50Aug 20, 2025Updated last year
- High performance inference engine for diffusion models☆107Sep 5, 2025Updated last year
- Some funny cute/cuteDSL code snippets☆35Mar 2, 2026Updated 6 months ago
- ☆16Sep 12, 2023Updated 3 years ago
- ☆67Oct 25, 2025Updated 10 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- FlashTile is a CUDA Tile IR compiler that is compatible with NVIDIA's tileiras, targeting SM70 through SM121 NVIDIA GPUs.☆60Feb 6, 2026Updated 7 months ago
- [WIP] Better (FP8) attention for Hopper☆33Aug 21, 2026Updated last month
- ☆32Jul 2, 2025Updated last year
- NVSHMEM‑Tutorial: Build a DeepEP‑like GPU Buffer☆204Feb 11, 2026Updated 7 months ago
- A Triton JIT runtime and ffi provider in C++☆41Updated this week
- torch_quantizer is a out-of-box quantization tool for PyTorch models on CUDA backend, specially optimized for Diffusion Models.☆25Mar 29, 2024Updated 2 years ago
- learn TensorRT from scratch🥰☆18Sep 29, 2024Updated last year
- FA4-based Relative Attention Kernel developed by TML and Colfax☆18Sep 11, 2026Updated last week
- high-performance linear attention kernel library built on TileLang☆706Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A high-performance RL training-inference weight synchronization framework, designed to enable second-level parameter updates from trainin…☆176Updated this week
- ☆34Jul 13, 2026Updated 2 months ago
- JAX bindings for the flash-attention3 kernels☆23Sep 9, 2026Updated 2 weeks ago
- Autonomous GPU kernel optimization system driven by AI agents.☆31Mar 29, 2026Updated 5 months ago
- ☆26May 30, 2025Updated last year
- [ICLR2025] Are Large Vision Language Models Good Game Players?☆13Mar 3, 2025Updated last year
- ☆21Mar 22, 2021Updated 5 years ago
- Awesome system papers for AI☆22Updated this week
- Combining Teacache with xDiT to Accelerate Visual Generation Models☆33Apr 21, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆19Mar 4, 2025Updated last year
- DiscreteTom's Blog Boilerplate.☆10Mar 6, 2023Updated 3 years ago
- ☆15Jun 21, 2023Updated 3 years ago
- Unified tracing profiler and visualizer, one timeline from CPU to GPU/HES to in-kernel zones☆147Updated this week
- ☆14Sep 7, 2024Updated 2 years ago
- SHOWMe: Benchmarking Object-agnostic Hand-Object 3D Reconstruction (Dataset, Contains proposed top baseline reconstructions with estimate…☆22Dec 18, 2023Updated 2 years ago
- Official implementation of "Towards One-Step Causal Video Generation via Adversarial Self-Distillation" (arXiv 2025). A novel framework f…☆32Nov 4, 2025Updated 10 months ago
- [NeurIPS 2024] ENAT: Rethinking Spatial-temporal Interactions in Token-based Image Synthesis☆25Nov 28, 2024Updated last year
- ☆46Oct 15, 2025Updated 11 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆34Feb 3, 2025Updated last year
- Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.☆369Updated this week
- ☆54May 19, 2025Updated last year
- [ACL 2026 Main] Code for the paper "ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs"☆32Updated this week
- Personal knowledge library☆11Nov 9, 2017Updated 8 years ago
- Created a simple neural network using C++17 standard and the Eigen library that supports both forward and backward propagation.☆11Jul 27, 2024Updated 2 years ago
- This repository provides tutorial, which discusses running sample publisher and subscriber using multiple transports of point_cloud_trans…☆11Aug 26, 2026Updated 3 weeks ago