This repositories contains the reference implementation for the Sparse Delta Memory paper.More precisely, it contains the model definition as well as triton and cuda kernels for the Sparse Delta Memory layer.
☆40Jul 9, 2026Updated 2 months ago
Alternatives and similar repositories for sparse-delta-memory
Users that are interested in sparse-delta-memory are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DeltaProduct is a new linear recurrent neural network architecture that uses products of generalized Householder matrices as state-transi…☆18Oct 13, 2025Updated 11 months ago
- Code and data to explore neural scaling laws of xLSTM and Transformer models.☆24Apr 8, 2026Updated 5 months ago
- Code for Fast-weight Product Key Memory (FwPKM)☆26Mar 18, 2026Updated 6 months ago
- ☆87Aug 19, 2026Updated last month
- ☆19Dec 12, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆18Jun 4, 2026Updated 3 months ago
- [Tech Report] Expanded Hyper-Connections☆68Jul 21, 2026Updated 2 months ago
- The Newton-Muon optimizer☆33Jun 5, 2026Updated 3 months ago
- Attention variant with per-channel multiplicative decay☆50Jun 3, 2026Updated 4 months ago
- Implementation of Fast Weight Attention☆35Sep 17, 2026Updated 2 weeks ago
- Official repository for Parallax (Parameterized Local Linear Attention)☆69Jul 30, 2026Updated 2 months ago
- Flash Attention in 300-500 lines of CUDA/C++☆39Aug 22, 2025Updated last year
- Multi-modal, multi-task modeling of the mouse visual cortex☆19May 5, 2026Updated 4 months ago
- High-performance GPU kernels for LLM inference in OpenAI Triton. Fused RMSNorm, SwiGLU, INT8 GEMM with benchmarks and roofline analysis.☆45Jul 22, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆27Jul 13, 2026Updated 2 months ago
- 更纯粹、更高压缩率的Tokenizer in Rust☆14Sep 7, 2026Updated 3 weeks ago
- Aurora optimizer release☆157Jul 18, 2026Updated 2 months ago
- Implementation and explorations into DiscoRL, Discovering state-of-the-art reinforcement learning algorithms, David Silver's last work at…☆22Jun 13, 2026Updated 3 months ago
- Official Project Page for HLA: Higher-order Linear Attention (https://arxiv.org/abs/2510.27258)☆105Jun 15, 2026Updated 3 months ago
- Expanding linear RNN state-transition matrix eigenvalues to include negatives improves state-tracking tasks and language modeling without…☆22Mar 15, 2025Updated last year
- Official repository for "SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space"☆32May 7, 2026Updated 4 months ago
- Experiments on the impact of depth in transformers and SSMs.☆47Oct 23, 2025Updated 11 months ago
- XWikisCorpus, cross-lingual summarisation, multi-lingual summarisation, pre-trained language models, zero-shot and few-shot summarisation…☆10Nov 4, 2022Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- A Multi-Policy, Multi-Agent RL Training Framework☆32Jun 16, 2026Updated 3 months ago
- Lean formalizations for the paper "On the paucity of lattice triangles"☆19Sep 11, 2026Updated 3 weeks ago
- Stick-breaking attention☆64Jul 1, 2025Updated last year
- Code with CliqueFlowmer model for Optimal Computational Materials Discovery☆17Sep 6, 2026Updated 3 weeks ago
- ☆31Aug 10, 2026Updated last month
- Delta Attention Residuals - supplementary code and pretrained models☆43Sep 26, 2026Updated last week
- 🔥 A minimal training framework for scaling FLA models☆420Apr 22, 2026Updated 5 months ago
- Code for ICLR 23: Text Summarization with Oracle Expectation☆13Sep 27, 2022Updated 4 years ago
- Efficient PScan implementation in PyTorch☆17Jan 2, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- This is the official code implementation of Bongard-OpenWorld (ICLR 2024).☆14Jan 6, 2025Updated last year
- ☆13Aug 19, 2024Updated 2 years ago
- mHC-lite: You Don’t Need 20 Sinkhorn-Knopp Iterations☆95Jan 12, 2026Updated 8 months ago
- An official implementation for the EMNLP 2023 Findings paper "Prompt-Based Editing for Text Style Transfer"☆13Dec 9, 2023Updated 2 years ago
- Fast Punctuation Restoration using Transformer Models for Vietnamese☆11Jun 10, 2022Updated 4 years ago
- Assist Non-native Viewers: Multimodal Crosslingual Summarization for How2 Videos☆10Sep 2, 2024Updated 2 years ago
- A framework to train language models to learn invariant representations.☆14Jan 24, 2022Updated 4 years ago