Distributed IO-aware Attention algorithm
☆24Sep 24, 2025Updated 10 months ago
Alternatives and similar repositories for Burst-Attention
Users that are interested in Burst-Attention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Sequence-level 1F1B schedule for LLMs.☆19Jun 4, 2024Updated 2 years ago
- Sequence-level 1F1B schedule for LLMs.☆37Aug 26, 2025Updated 11 months ago
- Nsight Compute In Docker☆13Dec 21, 2023Updated 2 years ago
- LLM training technologies developed by kwai☆71Jun 30, 2026Updated 3 weeks ago
- Official repository for DistFlashAttn: Distributed Memory-efficient Attention for Long-context LLMs Training☆223Aug 19, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Standalone Flash Attention v2 kernel without libtorch dependency☆113Sep 10, 2024Updated last year
- ☆41Jun 5, 2024Updated 2 years ago
- 🤖FFPA: Extends FA-2/3 via Split-D for large headdims, 1.5x~6×↑🎉 vs SDPA, up to 513~535 TFLOPS🎉 on NVIDIA H200.☆318Updated this week
- Zero Bubble Pipeline Parallelism☆464May 7, 2025Updated last year
- USP: Unified (a.k.a. Hybrid, 2D) Sequence Parallel Attention for Long Context Transformers Model Training and Inference☆683May 21, 2026Updated 2 months ago
- Ring attention implementation with flash attention☆1,038Sep 10, 2025Updated 10 months ago