Official PyTorch Implementation of Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
☆253May 25, 2026Updated 2 months ago
Alternatives and similar repositories for GatedDeltaNet-2
Users that are interested in GatedDeltaNet-2 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official repository for Parallax (Parameterized Local Linear Attention)☆68Jul 7, 2026Updated 3 weeks ago
- [ICLR 2025] Official PyTorch Implementation of Gated Delta Networks: Improving Mamba2 with Delta Rule☆634Mar 13, 2026Updated 4 months ago
- Delta Attention Residuals - supplementary code and pretrained models☆42May 20, 2026Updated 2 months ago
- Gecko Architecture☆16Jan 13, 2026Updated 6 months ago
- The Newton-Muon optimizer☆30Jun 5, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- high-performance linear attention kernel library built on TileLang☆619Updated this week
- 🚀 Efficient implementations for emerging model architectures☆5,473Updated this week
- FlashKDA: high-performance Kimi Delta Attention kernels☆982Updated this week
- Official Codebase: LT2: Linear-Time Looped Transformers.☆49Updated this week
- ☆284Jun 6, 2025Updated last year
- CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.☆535Jul 23, 2026Updated last week
- Attention variant with per-channel multiplicative decay☆48Jun 3, 2026Updated last month
- ☆22Jul 10, 2026Updated 2 weeks ago
- Code for Fast-weight Product Key Memory (FwPKM)☆19Mar 18, 2026Updated 4 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆1,540Nov 17, 2025Updated 8 months ago
- Official code for HiLS-Attention☆128Updated this week
- 🔥 A minimal training framework for scaling FLA models☆406Apr 22, 2026Updated 3 months ago
- [Tech Report] Expanded Hyper-Connections☆55Jul 21, 2026Updated last week
- implementations and experimentation on mHC by deepseek - https://arxiv.org/abs/2512.24880☆369Feb 17, 2026Updated 5 months ago
- Code for MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization☆22Feb 18, 2026Updated 5 months ago
- Cuda kernels for leveraging LLM sparsity to improve throughput and decrease the memory requirements during inference and training.☆255Jun 29, 2026Updated last month
- ☆385Updated this week
- LM engine is a library for pretraining/finetuning LLMs☆184Jul 23, 2026Updated last week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆15Updated this week
- Attempt to make multiple residual streams from Bytedance's Hyper-Connections paper accessible to the public☆187May 13, 2026Updated 2 months ago
- ♡☆22Updated this week
- DeltaProduct is a new linear recurrent neural network architecture that uses products of generalized Householder matrices as state-transi…☆15Oct 13, 2025Updated 9 months ago
- Code for accepted paper at ICLR 2026☆15May 19, 2026Updated 2 months ago
- Official JAX implementation of End-to-End Test-Time Training for Long Context☆627Feb 15, 2026Updated 5 months ago
- ☆51May 20, 2025Updated last year
- Fast Polar Decomposition for Muon☆169Jul 2, 2026Updated 3 weeks ago
- ☆91Jul 23, 2026Updated last week
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆139Feb 4, 2026Updated 5 months ago
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆34May 26, 2026Updated 2 months ago
- Triton kernels for dynamic causal short convolutions.☆24Jun 4, 2026Updated last month
- AdaSplash: Adaptive Sparse Flash Attention (aka Flash Entmax Attention)☆46May 20, 2026Updated 2 months ago
- The official code of "Mano: Restriking Manifold Optimization for LLM Training".☆25Jun 1, 2026Updated last month
- TypeScript SDK for building AI agents with automatic, scoped, persistent memory. `agent-memory` wraps model calls with memory recall and…☆24Jun 15, 2026Updated last month
- CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs☆235Updated this week