Linear Attention for Efficient Bidirectional Sequence Modeling
☆18May 13, 2025Updated last year
Alternatives and similar repositories for LION
Users that are interested in LION are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DeltaProduct is a new linear recurrent neural network architecture that uses products of generalized Householder matrices as state-transi…☆18Oct 13, 2025Updated 11 months ago
- PathPiece tokenizer☆14Nov 10, 2024Updated last year
- ☆16May 14, 2024Updated 2 years ago
- [NeurIPS 2024] Image Understanding Makes for A Good Tokenizer for Image Generation☆21Dec 17, 2024Updated last year
- High Performance Int8 GEMM Kernels for SM80 and later GPUs.☆24Mar 11, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A curated list of resources related to linear attention mechanisms.☆18Mar 16, 2025Updated last year
- This project contains code for the paper titled "SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentia…☆29Feb 21, 2024Updated 2 years ago
- Subset-Norm and Subset-Momentum. This repo is built on top of https://github.com/jiaweizzhao/GaLore.☆19Jul 9, 2025Updated last year
- ☆34Oct 22, 2024Updated last year
- Code for the Avey-B paper (https://arxiv.org/abs/2602.15814)☆33Feb 21, 2026Updated 7 months ago
- Training and evaluation code for the paper "Headless Language Models: Learning without Predicting with Contrastive Weight Tying" (https:/…☆30Apr 17, 2024Updated 2 years ago
- Official code for the NeurIPS25 paper "RAT: Bridging RNN Efficiencyand Attention Accuracy in Language Modeling" (https://arxiv.org/abs/25…☆26Aug 31, 2026Updated last month
- [ICML'25] "Rethinking Addressing in Language Models via Contextualized Equivariant Positional Encoding" by Jiajun Zhu, Peihao Wang, Ruisi…☆15Jun 6, 2025Updated last year
- ☆13Apr 23, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆19Nov 5, 2025Updated 11 months ago
- Attention variant with per-channel multiplicative decay☆50Jun 3, 2026Updated 4 months ago
- Casande-RL☆11May 9, 2023Updated 3 years ago
- RWKV-X is a Linear Complexity Hybrid Language Model based on the RWKV architecture, integrating Sparse Attention to improve the model's l…☆60Mar 31, 2026Updated 6 months ago
- BertSNR: an interpretable deep learning framework for single nucleotide resolution identification of transcription factor binding sites b…☆13May 27, 2025Updated last year
- [ICML-2025] We introduce Lie group Relative position Encodings (LieRE) that goes beyond RoPE in supporting n-dimensional inputs.☆14Aug 8, 2025Updated last year
- [CVPR 2023] Better “CMOS” Produces Clearer Images: Learning Space-Variant Blur Estimation for Blind Image Super-Resolution☆11Sep 14, 2026Updated 3 weeks ago
- HyPe: Better Pre-trained Language Model Fine-tuning with Hidden Representation Perturbation [ACL 2023]☆14Jul 11, 2023Updated 3 years ago
- Repository for the NeurIPS 2023 paper "Beyond Confidence: Reliable Models Should Also Consider Atypicality"☆13Apr 21, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [𝗜𝗖𝗠𝗟 𝟮𝟬𝟮𝟲] Dispersion loss counteracts embedding condensation and improves generalization in small language models☆21May 21, 2026Updated 4 months ago
- ☆12Dec 9, 2022Updated 3 years ago
- HGRN2: Gated Linear RNNs with State Expansion☆59Aug 20, 2024Updated 2 years ago
- ☆17Feb 14, 2025Updated last year
- OCDet: Object Center Detection via Bounding Box-Aware Heatmap Prediction on Edge Devices with NPUs☆12Jul 8, 2026Updated 3 months ago
- [CVPR2023] Practical Network Acceleration with Tiny Sets☆13Jul 28, 2023Updated 3 years ago
- [CIKM 2024] Do We Really Need Graph Convolution During Training? Light Post-Training Graph-ODE for Efficient Recommendation☆14Aug 11, 2024Updated 2 years ago
- ☆10Dec 17, 2020Updated 5 years ago
- This is the official repository for the paper "Laplacian Features for Learning with Hyperbolic Space"☆14Aug 8, 2022Updated 4 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A python package for DICOM to NifTi and NifTi to DICOM-SEG and GSPS conversion☆12Sep 25, 2023Updated 3 years ago
- [ICML 2024] VQDNA: Unleashing the Power of Vector Quantization for Multi-Species Genomic Sequence Modeling☆10Sep 22, 2024Updated 2 years ago
- ☆10Oct 2, 2024Updated 2 years ago
- A heterogeneous graph automatic meta-path learning method for drug-target interaction prediction☆14Aug 29, 2023Updated 3 years ago
- ☆13Jun 16, 2021Updated 5 years ago
- mPLM-Sim: Better Cross-Lingual Similarity and Transfer in Multilingual Pretrained Language Models☆11Jan 19, 2024Updated 2 years ago
- Provides a convenience Julia macro to extract fields from composite types☆15Feb 14, 2020Updated 6 years ago