Language modeling with linear-cost context
☆119Sep 25, 2025Updated 10 months ago
Alternatives and similar repositories for retention
Users that are interested in retention are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ROSA+: RWKV's ROSA implementation with fallback statistical predictor☆36Oct 13, 2025Updated 9 months ago
- A comprehensive AI & ML project portfolio from the University of Texas at Austin PG Program, demonstrating real-world data science and ma…☆17Jan 25, 2026Updated 6 months ago
- Marketplace ML experiment - training without backprop☆28Sep 9, 2025Updated 11 months ago
- ☆70Aug 3, 2026Updated last week
- https://x.com/BlinkDL_AI/status/1884768989743882276☆28May 4, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention (NeurIPS'25 Spotlight)☆26Feb 22, 2026Updated 5 months ago
- Training framework with a goal to explore the frontier of sample efficiency of small language models☆101Jan 25, 2026Updated 6 months ago
- A 20M RWKV v6 can do nonogram☆13Oct 18, 2024Updated last year
- ☆69Mar 21, 2025Updated last year
- MoE training for Me and You and maybe other people☆396Mar 15, 2026Updated 4 months ago
- ☆50Jul 3, 2026Updated last month
- Ludic – an LLM-RL library for the era of experience☆68Updated this week
- Optimized primitives for collective multi-GPU communication☆11May 8, 2024Updated 2 years ago
- 🔥 A minimal training framework for scaling FLA models☆410Apr 22, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Reference implementation of models from Nyonic Model Factory☆12May 13, 2024Updated 2 years ago
- Code repository for "RL Grokking Recipe: How RL Unlocks and Transfers New Algorithms in LLMs""☆35Oct 12, 2025Updated 9 months ago
- NeuroBLAST v3 architecture code☆37Jan 6, 2026Updated 7 months ago
- The official github repo for "Diffusion Language Models are Super Data Learners".☆228Nov 6, 2025Updated 9 months ago
- Mini Model Daemon☆13Nov 9, 2024Updated last year
- RWKV-X is a Linear Complexity Hybrid Language Model based on the RWKV architecture, integrating Sparse Attention to improve the model's l…☆59Mar 31, 2026Updated 4 months ago
- minimal Energy-based transformer☆44Dec 11, 2025Updated 8 months ago
- A framework for optimizing DSPy programs with RL☆341Jan 12, 2026Updated 6 months ago
- ☆39Aug 1, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆16Jun 15, 2026Updated last month
- Meta-Reinforcement Learning with Self-Reflection☆33Mar 26, 2026Updated 4 months ago
- ☆93Aug 18, 2024Updated last year
- Residual vector quantization for KV cache compression in large language model☆12Oct 22, 2024Updated last year
- Experiments on GPT-3's ability to fit numerical models in-context.☆14Aug 11, 2022Updated 4 years ago
- Scalable and Stable Parallelization of Nonlinear RNNS☆33Jun 28, 2026Updated last month
- MultiscaleGraphSignalTransforms.jl is a collection of software tools written in the Julia programming language for graph signal processin…☆12Mar 15, 2026Updated 4 months ago
- MLX implementation of Hierarchical Reasoning Model (HRM) - Adaptive computation for complex reasoning tasks☆29Aug 27, 2025Updated 11 months ago
- 100M tokens. Infinite compute. Lowest val loss wins.☆523Jul 3, 2026Updated last month
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Learning Accurate Decision Trees with Bandit Feedback via Quantized Gradient Descent☆16Sep 8, 2022Updated 3 years ago
- Research work aimed at addressing the problem of modeling infinite-length context☆52Dec 18, 2025Updated 7 months ago
- A Triton Kernel for incorporating Bi-Directionality in Mamba2☆83Dec 18, 2024Updated last year
- [NeurIPS 2025] Reinforcement Learning for Reasoning in Large Language Models with One Training Example☆444Mar 11, 2026Updated 5 months ago
- Chain-of-thought 방식을 활용하여 llama2를 fine-tuning☆10Nov 18, 2023Updated 2 years ago
- Direct Preference Optimization for RWKV, aiming for RWKV-5 and 6.☆11Mar 1, 2024Updated 2 years ago
- Official Chinese documentation for RWKV | RWKV官方中文文档☆15Jun 10, 2026Updated 2 months ago