Official repository for DistFlashAttn: Distributed Memory-efficient Attention for Long-context LLMs Training
☆222Aug 19, 2024Updated 2 years ago
Alternatives and similar repositories for LightSeq
Users that are interested in LightSeq are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Python package for rematerialization-aware gradient checkpointing☆27Oct 31, 2023Updated 2 years ago
- Ring attention implementation with flash attention☆1,050Sep 10, 2025Updated 11 months ago
- USP: Unified (a.k.a. Hybrid, 2D) Sequence Parallel Attention for Long Context Transformers Model Training and Inference☆690May 21, 2026Updated 3 months ago
- Large Context Attention☆772Oct 13, 2025Updated 10 months ago
- Zero Bubble Pipeline Parallelism☆466May 7, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆49Nov 10, 2023Updated 2 years ago
- Reference implementation of "Softmax Attention with Constant Cost per Token" (Heinsen, 2024)☆25Jun 6, 2024Updated 2 years ago
- PyTorch bindings for CUTLASS grouped GEMM.☆192Apr 8, 2026Updated 4 months ago
- Memory optimization and training recipes to extrapolate language models' context length to 1 million tokens, with minimal hardware.☆761Sep 27, 2024Updated last year
- Distributed Compiler and Optimized Parallel Kernels☆1,532Aug 12, 2026Updated 3 weeks ago
- Code for the paper "Rethinking Benchmark and Contamination for Language Models with Rephrased Samples"☆325Dec 20, 2023Updated 2 years ago
- Linear Attention Sequence Parallelism (LASP)☆87Jun 4, 2024Updated 2 years ago
- FlexAttention w/ FlashAttention3 Support☆27Oct 5, 2024Updated last year
- Implementation of 💍 Ring Attention, from Liu et al. at Berkeley AI, in Pytorch☆545May 16, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official repository for LongChat and LongEval☆535May 24, 2024Updated 2 years ago
- [ICML'24] Data and code for our paper "Training-Free Long-Context Scaling of Large Language Models"