Language modeling with linear-cost context
☆122Sep 25, 2025Updated 11 months ago
Alternatives and similar repositories for retention
Users that are interested in retention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Attention Kernels for Symmetric Power Transformers☆131Sep 25, 2025Updated 11 months ago
- ROSA+: RWKV's ROSA implementation with fallback statistical predictor☆36Oct 13, 2025Updated 11 months ago
- A comprehensive AI & ML project portfolio from the University of Texas at Austin PG Program, demonstrating real-world data science and ma…☆18Jan 25, 2026Updated 7 months ago
- Marketplace ML experiment - training without backprop☆28Sep 9, 2025Updated last year
- ☆81Sep 11, 2026Updated last week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- https://x.com/BlinkDL_AI/status/1884768989743882276☆28May 4, 2025Updated last year
- Training framework with a goal to explore the frontier of sample efficiency of small language models☆101Jan 25, 2026Updated 7 months ago
- ☆30Feb 27, 2024Updated 2 years ago
- A 20M RWKV v6 can do nonogram☆13Oct 18, 2024Updated last year
- ☆69Mar 21, 2025Updated last year
- MoE training for Me and You and maybe other people☆400Mar 15, 2026Updated 6 months ago
- ☆51Jul 3, 2026Updated 2 months ago
- Ludic – an LLM-RL library for the era of experience☆67Aug 9, 2026Updated last month
- Optimized primitives for collective multi-GPU communication☆11May 8, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆52Feb 17, 2025Updated last year
- ☆81Feb 18, 2026Updated 7 months ago
- 🔥 A minimal training framework for scaling FLA models☆416Apr 22, 2026Updated 5 months ago
- Reference implementation of models from Nyonic Model Factory☆12May 13, 2024Updated 2 years ago
- Code repository for "RL Grokking Recipe: How RL Unlocks and Transfers New Algorithms in LLMs""☆35Oct 12, 2025Updated 11 months ago
- Official implementation of "Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers"☆177Jan 30, 2025Updated last year
- NeuroBLAST v3 architecture code☆37Jan 6, 2026Updated 8 months ago
- The official github repo for "Diffusion Language Models are Super Data Learners".☆228Nov 6, 2025Updated 10 months ago
- Mini Model Daemon☆13Nov 9, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- minimal Energy-based transformer☆44Dec 11, 2025Updated 9 months ago
- RWKV-X is a Linear Complexity Hybrid Language Model based on the RWKV architecture, integrating Sparse Attention to improve the model's l…☆60Mar 31, 2026Updated 5 months ago
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence☆70Nov 11, 2025Updated 10 months ago
- A framework for optimizing DSPy programs with RL☆339Jan 12, 2026Updated 8 months ago
- ☆41Aug 1, 2025Updated last year
- Two implementations of ZeRO-1 optimizer sharding in JAX☆14Jun 11, 2023Updated 3 years ago
- ☆198Dec 18, 2025Updated 9 months ago
- Meta-Reinforcement Learning with Self-Reflection☆36Mar 26, 2026Updated 5 months ago
- ☆93Aug 18, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Lightweight piece tokenization library☆12Aug 24, 2026Updated 3 weeks ago
- Official code and dataset repository of KoBBQ (TACL 2024)☆23May 13, 2024Updated 2 years ago
- ☆27May 14, 2026Updated 4 months ago
- FlexAttention w/ FlashAttention3 Support☆27Oct 5, 2024Updated last year
- Residual vector quantization for KV cache compression in large language model☆12Oct 22, 2024Updated last year
- Scalable and Stable Parallelization of Nonlinear RNNS☆33Jun 28, 2026Updated 2 months ago
- Agentic RL Training at Scale☆2,068Updated this week