This repositories contains the reference implementation for the Sparse Delta Memory paper.More precisely, it contains the model definition as well as triton and cuda kernels for the Sparse Delta Memory layer.
☆32Jul 9, 2026Updated 3 weeks ago
Alternatives and similar repositories for sparse-delta-memory
Users that are interested in sparse-delta-memory are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DeltaProduct is a new linear recurrent neural network architecture that uses products of generalized Householder matrices as state-transi…☆15Oct 13, 2025Updated 9 months ago
- Code for Fast-weight Product Key Memory (FwPKM)☆20Mar 18, 2026Updated 4 months ago
- ☆78May 29, 2026Updated 2 months ago
- ☆19Dec 12, 2023Updated 2 years ago
- ☆16Jun 4, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [Tech Report] Expanded Hyper-Connections☆57Jul 21, 2026Updated 2 weeks ago
- Parallel evolution pipeline for Karpathy's autoresearch on SageMaker Spot Training (H100). 10x faster with HUGI pattern.☆23Apr 17, 2026Updated 3 months ago
- Attention variant with per-channel multiplicative decay☆48Jun 3, 2026Updated 2 months ago
- Official repository for Parallax (Parameterized Local Linear Attention)☆68Updated this week
- Flash-Linear-Attention models beyond language☆21Aug 28, 2025Updated 11 months ago
- Flash Attention in 300-500 lines of CUDA/C++☆39Aug 22, 2025Updated 11 months ago
- Multi-modal, multi-task modeling of the mouse visual cortex☆16May 5, 2026Updated 2 months ago
- High-performance GPU kernels for LLM inference in OpenAI Triton. Fused RMSNorm, SwiGLU, INT8 GEMM with benchmarks and roofline analysis.☆34Jul 22, 2026Updated last week
- ☆25Jul 13, 2026Updated 3 weeks ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Aurora optimizer release☆151Jul 18, 2026Updated 2 weeks ago
- Implementation and explorations into DiscoRL, Discovering state-of-the-art reinforcement learning algorithms, David Silver's last work at…☆21Jun 13, 2026Updated last month
- Official Project Page for HLA: Higher-order Linear Attention (https://arxiv.org/abs/2510.27258)☆102Jun 15, 2026Updated last month
- Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free…☆24Jul 26, 2026Updated last week
- Expanding linear RNN state-transition matrix eigenvalues to include negatives improves state-tracking tasks and language modeling without…☆22Mar 15, 2025Updated last year
- Official repository for "SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space"☆27May 7, 2026Updated 2 months ago
- Experiments on the impact of depth in transformers and SSMs.☆45Oct 23, 2025Updated 9 months ago
- XWikisCorpus, cross-lingual summarisation, multi-lingual summarisation, pre-trained language models, zero-shot and few-shot summarisation…☆10Nov 4, 2022Updated 3 years ago
- A Multi-Policy, Multi-Agent RL Training Framework☆32Jun 16, 2026Updated last month
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- STABILIZING GRADIENTS FOR DEEP NEURAL NETWORKS VIA EFFICIENT SVD PARAMETERIZATION☆16Jun 5, 2018Updated 8 years ago
- ☆25Dec 29, 2018Updated 7 years ago
- Lean formalizations for the paper "On the paucity of lattice triangles"☆19Mar 26, 2026Updated 4 months ago
- ☆48Dec 13, 2025Updated 7 months ago
- Code with CliqueFlowmer model for Optimal Computational Materials Discovery☆17Apr 21, 2026Updated 3 months ago
- ☆30Apr 28, 2026Updated 3 months ago
- Delta Attention Residuals - supplementary code and pretrained models☆42May 20, 2026Updated 2 months ago
- ☆18Apr 17, 2026Updated 3 months ago
- 🔥 A minimal training framework for scaling FLA models☆409Apr 22, 2026Updated 3 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆13Sep 27, 2022Updated 3 years ago
- Efficient PScan implementation in PyTorch☆17Jan 2, 2024Updated 2 years ago
- This is the official code implementation of Bongard-OpenWorld (ICLR 2024).☆14Jan 6, 2025Updated last year
- ☆13Aug 19, 2024Updated last year
- mHC-lite: You Don’t Need 20 Sinkhorn-Knopp Iterations☆91Jan 12, 2026Updated 6 months ago
- An official implementation for the EMNLP 2023 Findings paper "Prompt-Based Editing for Text Style Transfer"☆13Dec 9, 2023Updated 2 years ago
- Assist Non-native Viewers: Multimodal Crosslingual Summarization for How2 Videos☆10Sep 2, 2024Updated last year