☆54Dec 12, 2023Updated 2 years ago
Alternatives and similar repositories for looped_transformer
Users that are interested in looped_transformer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official implementation of the transformer (TF) architecture suggested in a paper entitled "Looped Transformers as Programmable Computers…☆43Apr 8, 2023Updated 3 years ago
- ☆20Oct 25, 2022Updated 3 years ago
- Generative Equilibrium Transformer☆28Nov 11, 2023Updated 2 years ago
- Combining SOAP and MUON☆25Feb 11, 2025Updated last year
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence☆69Nov 11, 2025Updated 9 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆35May 26, 2026Updated 2 months ago
- ☆15Mar 20, 2025Updated last year
- Complexity Based Prompting for Multi-Step Reasoning☆17Mar 10, 2023Updated 3 years ago
- [ICML 2025] Unlearning in Diffusion Models using Sparse Autoencoders☆63Oct 16, 2025Updated 10 months ago
- Gecko Architecture☆18Jan 13, 2026Updated 7 months ago
- Code accompanying the paper "A contrastive rule for meta-learning"☆13Oct 31, 2024Updated last year
- Code for "Inducer-tuning: Connecting Prefix-tuning and Adapter-tuning" (EMNLP 2022) and "Empowering Parameter-Efficient Transfer Learning…☆11Feb 6, 2023Updated 3 years ago
- ☆21Mar 1, 2023Updated 3 years ago
- Official implementation of the paper "Linear Transformers with Learnable Kernel Functions are Better In-Context Models"☆169Jan 16, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆13Sep 18, 2024Updated last year
- ☆93Aug 18, 2024Updated 2 years ago
- Stick-breaking attention☆63Jul 1, 2025Updated last year
- Supervised Training of Conditional Monge Maps☆20Oct 30, 2023Updated 2 years ago
- ☆45Apr 30, 2018Updated 8 years ago
- Efficient PScan implementation in PyTorch☆17Jan 2, 2024Updated 2 years ago
- Residual vector quantization for KV cache compression in large language model☆12Oct 22, 2024Updated last year
- Research Papers on Efficient Neural Fields from EffL Group☆16Apr 21, 2025Updated last year
- An approximate implementation of the OpenAI paper - An Empirical Model of Large-Batch Training for MNIST☆11Nov 19, 2022Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆14Mar 2, 2025Updated last year
- 🧮 Algebraic Positional Encodings.☆21Jun 5, 2026Updated 2 months ago
- ☆67May 18, 2023Updated 3 years ago
- this repository contains the code and experimental setup for the cikm 2025 paper “llm4es: learning user embeddings from event sequences v…☆18Apr 7, 2026Updated 4 months ago
- ☆18May 25, 2023Updated 3 years ago
- The official github repo for "Training Optimal Large Diffusion Language Models", the first-ever large-scale diffusion language models sca…☆46Nov 6, 2025Updated 9 months ago
- Learning Accurate Decision Trees with Bandit Feedback via Quantized Gradient Descent☆16Sep 8, 2022Updated 3 years ago
- Official Code Repository for the paper "Key-value memory in the brain"☆33Feb 25, 2025Updated last year
- The is the official implementation of "Lyra: Orchestrating Dual Correction in Automated Theorem Proving"☆15Jul 2, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- our submission for the microsoft membership inference competion at SaTML 2023☆15Apr 5, 2023Updated 3 years ago
- [NeurIPS 2024 D&B Track] UnlearnCanvas: A Stylized Image Dataset to Benchmark Machine Unlearning for Diffusion Models by Yihua Zhang, Cho…☆88Nov 11, 2024Updated last year
- Learning Universal Predictors☆85Aug 1, 2024Updated 2 years ago
- Attempt to make multiple residual streams from Bytedance's Hyper-Connections paper accessible to the public☆188May 13, 2026Updated 3 months ago
- ☆12Nov 22, 2024Updated last year
- ☆14Mar 7, 2024Updated 2 years ago
- MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic Approximation☆18Sep 2, 2024Updated last year