Language modeling with linear-cost context
☆121Sep 25, 2025Updated 11 months ago
Alternatives and similar repositories for retention
Users that are interested in retention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A framework for clean, testable, and high-performance CUDA kernels.☆22Sep 23, 2025Updated 11 months ago
- Attention Kernels for Symmetric Power Transformers☆131Sep 25, 2025Updated 11 months ago
- ROSA+: RWKV's ROSA implementation with fallback statistical predictor☆36Oct 13, 2025Updated 10 months ago
- A comprehensive AI & ML project portfolio from the University of Texas at Austin PG Program, demonstrating real-world data science and ma…☆18Jan 25, 2026Updated 7 months ago
- Marketplace ML experiment - training without backprop☆28Sep 9, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆80Aug 14, 2026Updated 2 weeks ago
- https://x.com/BlinkDL_AI/status/1884768989743882276☆28May 4, 2025Updated last year
- MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention (NeurIPS'25 Spotlight)☆26Feb 22, 2026Updated 6 months ago
- Training framework with a goal to explore the frontier of sample efficiency of small language models☆101Jan 25, 2026Updated 7 months ago
- A 20M RWKV v6 can do nonogram☆13Oct 18, 2024Updated last year
- ☆69Mar 21, 2025Updated last year
- MoE training for Me and You and maybe other people☆398Mar 15, 2026Updated 5 months ago
- ☆51Jul 3, 2026Updated last month
- Ludic – an LLM-RL library for the era of experience☆67Aug 9, 2026Updated 3 weeks ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Optimized primitives for collective multi-GPU communication☆11May 8, 2024Updated 2 years ago
- ☆52Feb 17, 2025Updated last year
- Official implementation of "Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers"☆176Jan 30, 2025Updated last year
- Reference implementation of models from Nyonic Model Factory☆12May 13, 2024Updated 2 years ago
- Code repository for "RL Grokking Recipe: How RL Unlocks and Transfers New Algorithms in LLMs""☆35Oct 12, 2025Updated 10 months ago
- NeuroBLAST v3 architecture code☆37Jan 6, 2026Updated 7 months ago
- The official github repo for "Diffusion Language Models are Super Data Learners".☆228Nov 6, 2025Updated 9 months ago
- Mini Model Daemon☆13Nov 9, 2024Updated last year
- RWKV-X is a Linear Complexity Hybrid Language Model based on the RWKV architecture, integrating Sparse Attention to improve the model's l…☆60Mar 31, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence☆69Nov 11, 2025Updated 9 months ago
- minimal Energy-based transformer☆44Dec 11, 2025Updated 8 months ago
- A framework for optimizing DSPy programs with RL☆339Jan 12, 2026Updated 7 months ago
- ☆40Aug 1, 2025Updated last year
- ☆17Jun 15, 2026Updated 2 months ago
- Two implementations of ZeRO-1 optimizer sharding in JAX☆14Jun 11, 2023Updated 3 years ago
- Meta-learning inductive biases in the form of useful conserved quantities.☆41Nov 19, 2022Updated 3 years ago
- ☆198Dec 18, 2025Updated 8 months ago
- Meta-Reinforcement Learning with Self-Reflection☆34Mar 26, 2026Updated 5 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆93Aug 18, 2024Updated 2 years ago
- Lightweight piece tokenization library☆12Aug 24, 2026Updated last week
- ☆41Apr 30, 2025Updated last year
- ☆26May 14, 2026Updated 3 months ago
- FlexAttention w/ FlashAttention3 Support☆27Oct 5, 2024Updated last year
- Residual vector quantization for KV cache compression in large language model☆12Oct 22, 2024Updated last year
- Experiments on GPT-3's ability to fit numerical models in-context.☆14Aug 11, 2022Updated 4 years ago