Official Codebase: LT2: Linear-Time Looped Transformers.
☆63Jul 27, 2026Updated 2 months ago
Alternatives and similar repositories for LT2
Users that are interested in LT2 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026] Spherical Steering: Geometry-Aware Activation Rotation for Language Models☆23May 19, 2026Updated 4 months ago
- [ICML 2026] Code for Equilibrium Reasoners: learning attractor dynamics for scalable reasoning☆57Sep 14, 2026Updated 2 weeks ago
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆35May 26, 2026Updated 4 months ago
- The best ChatGPT that $100 can buy.☆59Updated this week
- A curated list of papers and selected technical blogs on Loop Models.☆431Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆87Aug 19, 2026Updated last month
- [Tech Report] Expanded Hyper-Connections☆68Jul 21, 2026Updated 2 months ago
- ☆71Dec 6, 2023Updated 2 years ago
- Official repository for Parallax (Parameterized Local Linear Attention)☆69Jul 30, 2026Updated 2 months ago
- The official implementation of HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization☆19Mar 7, 2025Updated last year
- FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale☆80Updated this week
- Official PyTorch Implementation of Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention☆324Aug 29, 2026Updated last month
- ☆31May 20, 2026Updated 4 months ago
- The Newton-Muon optimizer☆33Jun 5, 2026Updated 3 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Triton kernels for dynamic causal short convolutions.☆30Jun 4, 2026Updated 3 months ago
- Cuda kernels for leveraging LLM sparsity to improve throughput and decrease the memory requirements during inference and training.☆263Jun 29, 2026Updated 3 months ago
- ☆26Dec 18, 2024Updated last year
- LoopFormer is an elastic-depth looped Transformer trained on variable-length trajectories, using time/step-size conditioning and a shortc…☆48Mar 28, 2026Updated 6 months ago
- Official Repo for Error-Free Linear Attention is a Free Lunch: Exact Solution from Continuous-Time Dynamics☆77Sep 25, 2026Updated last week
- ☆16Jul 16, 2024Updated 2 years ago
- ☆37Aug 7, 2025Updated last year
- [NeurIPS 2025] Official Pytorch Implementation of "The Curse of Depth in Large Language Models" by Wenfang Sun, Xinyuan Song, Pengxiang L…☆74Mar 3, 2026Updated 7 months ago
- Code and data to explore neural scaling laws of xLSTM and Transformer models.☆24Apr 8, 2026Updated 5 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆34Feb 1, 2026Updated 8 months ago
- A tiny FP8 multiplication unit written in Verilog. TinyTapeout 2 submission.☆14Nov 23, 2022Updated 3 years ago
- Official codebase for "Next-Latent Prediction Transformers Learn Compact World Models"☆195Jun 12, 2026Updated 3 months ago
- Official implementation of NeurIPS'24 paper "Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles"☆16Mar 26, 2025Updated last year
- ☆15May 2, 2026Updated 5 months ago
- [ICLR 2025] Official Pytorch Implementation of "Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN" by Pengxia…☆30Jul 24, 2025Updated last year
- ☆53May 20, 2025Updated last year
- ☆17Jun 15, 2026Updated 3 months ago
- DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation☆319Feb 18, 2026Updated 7 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Triton kernels and PyTorch ops for Block Attention Residuals (AttnRes)☆89May 29, 2026Updated 4 months ago
- Kwai Summary Attention☆60Aug 31, 2026Updated last month
- The implementation for ThreadWeaver Adaptive Threading for Efficient Parallel Reasoning in Language Models☆67Apr 8, 2026Updated 5 months ago
- Code for NeurIPS 2023 paper "Non-autoregressive Machine Translation with Probabilistic Context-free Grammar".☆12Jan 4, 2024Updated 2 years ago
- The official repo for "OpenMoE 2: Sparse Diffusion Language Models".☆58Dec 28, 2025Updated 9 months ago
- SR²AM: Efficient Agentic Reasoning Through Self-Regulated Simulative Planning☆21May 22, 2026Updated 4 months ago
- ☆23Mar 31, 2026Updated 6 months ago