DeepSeek Native Sparse Attention pytorch implementation
β119Dec 17, 2025Updated 7 months ago
Alternatives and similar repositories for NSA-pytorch
Users that are interested in NSA-pytorch are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β48Dec 13, 2025Updated 7 months ago
- π³ Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"β1,012Feb 5, 2026Updated 5 months ago
- PyTorch DeepSeek Sparse Attention (DSA) training & inferenceβ23Oct 14, 2025Updated 9 months ago
- Implementation of the sparse attention pattern proposed by the Deepseek team in their "Native Sparse Attention" paperβ809Aug 15, 2025Updated 11 months ago
- qwen-nsaβ87Oct 14, 2025Updated 9 months ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- β102Feb 11, 2026Updated 5 months ago
- An efficient implementation of the NSA (Native Sparse Attention) kernelβ133Jun 24, 2025Updated last year
- BigBang-Proton is a LLM pretrained on cross-scale, cross-structure, cross-discipline real-world scientific tasks to construct a scientiβ¦β21Nov 8, 2025Updated 8 months ago
- Code for paper: [ICLR2025 Oral] FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inferenceβ170Oct 13, 2025Updated 9 months ago
- High performance RMSNorm Implement by using SM Core Storage(Registers and Shared Memory)β30Jan 22, 2026Updated 6 months ago
- Efficient triton implementation of Native Sparse Attention.β284May 23, 2025Updated last year
- We introduce UltraLLaDA , a scaled variant of LLaDA-8B-Base that extends the context length up to 128K tokens with light-weight post-traiβ¦β15Oct 23, 2025Updated 9 months ago
- FlashTile is a CUDA Tile IR compiler that is compatible with NVIDIA's tileiras, targeting SM70 through SM121 NVIDIA GPUs.β61Feb 6, 2026Updated 5 months ago
- β251Nov 19, 2025Updated 8 months ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ASP-DAC 2025] LightCL: Compact Continual Learning with Low Memory Footprint For Edge Deviceβ16Apr 2, 2025Updated last year
- β37Aug 7, 2025Updated 11 months ago
- Some funny cute/cuteDSL code snippetsβ33Mar 2, 2026Updated 4 months ago
- β156Mar 4, 2025Updated last year
- analyse problems of AI with Math and Codeβ31Jul 28, 2025Updated 11 months ago
- A collection of memory efficient attention operators implemented in the Triton language.β300Updated this week
- [AAAI 2026 Oral] SAMCL: Empowering SAM to Continually Learn from Dynamic Domains with Extreme Storage Efficiencyβ21Mar 25, 2026Updated 3 months ago
- Visualize the Expert Parallelism Load Balancerβ19Mar 15, 2025Updated last year
- β45Oct 15, 2025Updated 9 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- β16Updated this week
- [NeurIPS 2024] Official implementation of paper "Perceiving Longer Sequences With Bi-Directional Cross-Attention Transformers"β22Mar 10, 2025Updated last year
- DeepSeek-V3.2-Exp DSA Warmup Lightning Indexer training operator based on tilelangβ47Nov 19, 2025Updated 8 months ago
- [ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Headsβ539Feb 10, 2025Updated last year
- A Survey of Efficient Attention Methods: Hardware-efficient, Sparse, Compact, and Linear Attentionβ304Dec 1, 2025Updated 7 months ago
- Design hardware-friendly model architectures and migrate existing LLMs with minimal performance lossβ491Jul 14, 2026Updated last week
- π Efficient implementations for emerging model architecturesβ5,394Updated this week
- The official repository for SkyLadder: Better and Faster Pretraining via Context Window Schedulingβ43Dec 29, 2025Updated 6 months ago
- pytorch implementation of DeepSeek Engramβ19Mar 24, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Cute layout visualizationβ44Jan 18, 2026Updated 6 months ago
- TileFusion is an experimental C++ macro kernel template library that elevates the abstraction level in CUDA C for tile processing.β115Jun 28, 2025Updated last year
- A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of β¦β329Jun 10, 2025Updated last year
- [ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.β1,017Feb 25, 2026Updated 4 months ago
- β135May 29, 2025Updated last year
- Efficient LLM Inference over Long Sequencesβ392Jun 25, 2025Updated last year
- Accelerating MoE with IO and Tile-aware Optimizationsβ732Jul 4, 2026Updated 2 weeks ago