Official Codebase: LT2: Linear-Time Looped Transformers.
☆50Jul 27, 2026Updated last week
Alternatives and similar repositories for LT2
Users that are interested in LT2 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026] Spherical Steering: Geometry-Aware Activation Rotation for Language Models☆18May 19, 2026Updated 2 months ago
- [ICML 2026] Code for Equilibrium Reasoners: learning attractor dynamics for scalable reasoning☆45Jun 1, 2026Updated 2 months ago
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆34May 26, 2026Updated 2 months ago
- A curated list of papers and selected technical blogs on Loop Models.☆248Updated this week
- ☆78May 29, 2026Updated 2 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆67Dec 6, 2023Updated 2 years ago
- [Tech Report] Expanded Hyper-Connections☆57Jul 21, 2026Updated 2 weeks ago
- Official repository for Parallax (Parameterized Local Linear Attention)☆68Updated this week
- FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale☆50May 30, 2026Updated 2 months ago
- Triton kernels for dynamic causal short convolutions.☆25Jun 4, 2026Updated last month
- Cuda kernels for leveraging LLM sparsity to improve throughput and decrease the memory requirements during inference and training.☆255Jun 29, 2026Updated last month
- ☆30Sep 16, 2023Updated 2 years ago
- LoopFormer is an elastic-depth looped Transformer trained on variable-length trajectories, using time/step-size conditioning and a shortc…☆30Mar 28, 2026Updated 4 months ago
- Official PyTorch Implementation of Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention☆257May 25, 2026Updated 2 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆25Dec 18, 2024Updated last year
- [ICML 2026] Public repository for fine-tuning Masked Diffusion Models toward provable self-correction.☆26Jul 5, 2026Updated 3 weeks ago
- ☆16Jul 16, 2024Updated 2 years ago
- Official codebase for "Next-Latent Prediction Transformers Learn Compact World Models"☆148Jun 12, 2026Updated last month
- Code and data to explore neural scaling laws of xLSTM and Transformer models.☆23Apr 8, 2026Updated 3 months ago
- Official implementation of NeurIPS'24 paper "Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles"☆16Mar 26, 2025Updated last year
- ☆15May 2, 2026Updated 3 months ago
- [ICLR 2025] Official Pytorch Implementation of "Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN" by Pengxia…☆30Jul 24, 2025Updated last year
- ☆16Jun 15, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation☆244Feb 18, 2026Updated 5 months ago
- Kwai Summary Attention☆59May 8, 2026Updated 2 months ago
- The implementation for ThreadWeaver Adaptive Threading for Efficient Parallel Reasoning in Language Models☆67Apr 8, 2026Updated 3 months ago
- The Newton-Muon optimizer☆29Jun 5, 2026Updated last month
- Code for NeurIPS 2023 paper "Non-autoregressive Machine Translation with Probabilistic Context-free Grammar".☆12Jan 4, 2024Updated 2 years ago
- SR²AM: Efficient Agentic Reasoning Through Self-Regulated Simulative Planning☆21May 22, 2026Updated 2 months ago
- The official repo for "OpenMoE 2: Sparse Diffusion Language Models".☆58Dec 28, 2025Updated 7 months ago
- Vstream - Video Analytics pipeline with Hardware based accelerations (dev - stage)☆10Feb 2, 2024Updated 2 years ago
- ☆17Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official Repo for Error-Free Linear Attention is a Free Lunch: Exact Solution from Continuous-Time Dynamics☆76Mar 26, 2026Updated 4 months ago
- Official Code for Paper "Think While You Generate: Discrete Diffusion with Planned Denoising" [ICLR 2025]☆85Apr 24, 2025Updated last year
- ☆17Jul 31, 2025Updated last year
- SRSA: Skill Retrieval and Adaptation for Robotic Assembly Tasks☆18Mar 25, 2026Updated 4 months ago
- [NeurIPS 2025] Official Pytorch Implementation of "The Curse of Depth in Large Language Models" by Wenfang Sun, Xinyuan Song, Pengxiang L…☆72Mar 3, 2026Updated 5 months ago
- A demonstrative example of running SGLang Diffusion with DP router☆17Mar 15, 2026Updated 4 months ago
- Evolution Pretraining Fully in Int Formats☆178Feb 25, 2026Updated 5 months ago