Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
☆69Nov 11, 2025Updated 9 months ago
Alternatives and similar repositories for retrofitting-recurrence
Users that are interested in retrofitting-recurrence are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆31May 21, 2026Updated 3 months ago
- Code for "DynaGuard: A Dynamic Guardrail Model With User-Defined Policies."☆23Nov 3, 2025Updated 10 months ago
- Gemstones: A Model Suite for Multi-Faceted Scaling Laws (NeurIPS 2025)☆35Sep 28, 2025Updated 11 months ago
- Official release of code for the paper RL is a hammer and LLMs are nails A simple RL approach to stronger prompt injection attacks☆53May 6, 2026Updated 3 months ago
- Code for the paper Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs☆70Jun 23, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- What do we learn from inverting CLIP models?☆58Mar 6, 2024Updated 2 years ago
- [NeurIPS 2024] Goldfish Loss: Mitigating Memorization in Generative LLMs☆98Nov 17, 2024Updated last year
- PyTorch Implementation of Zero-Shot Vision Encoder Grafting via LLM Surrogates [ICCV'25]☆54Jul 10, 2025Updated last year
- [ICML'26] Official implementation of paper "Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models"☆81Jul 17, 2026Updated last month
- Code to reproduce "Transformers Can Do Arithmetic with the Right Embeddings", McLeish et al (NeurIPS 2024)☆200May 28, 2024Updated 2 years ago
- ☆152Nov 22, 2025Updated 9 months ago
- Official repo for Detecting, Explaining, and Mitigating Memorization in Diffusion Models (ICLR 2024)☆80Apr 3, 2024Updated 2 years ago
- Stable Looped Models and their Scaling Laws☆174May 17, 2026Updated 3 months ago
- Code and data for paper "(How) do Language Models Track State?"☆28Mar 31, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Pretraining and inference code for a large-scale depth-recurrent language model☆920Dec 29, 2025Updated 8 months ago
- Learning from Mixed Rollouts: Logit Fusion as a Bridge Between Imitation and Exploration☆18Feb 24, 2026Updated 6 months ago
- Measuring the Signal to Noise Ratio in Language Model Evaluation☆31Aug 19, 2025Updated last year
- Official Code for "Baseline Defenses for Adversarial Attacks Against Aligned Language Models"☆34Oct 26, 2023Updated 2 years ago
- ☆18Oct 12, 2022Updated 3 years ago
- ACL 2026 & NAACL 2025: Bridging Retrieval and Inference through Evidence Fusion☆14Apr 9, 2026Updated 4 months ago
- ☆17Updated this week
- ☆116Jun 8, 2026Updated 2 months ago
- The Official Repository for "Bring Your Own Data! Self-Supervised Evaluation for Large Language Models"☆108Sep 23, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official implementation of the paper "Pretraining Language Models to Ponder in Continuous Space"☆27Jul 21, 2025Updated last year
- [COLM 2025: 1st Workshop on the Application of LLM Explainability to Reasoning and Planning] Latent Chain-of-Thought? Decoding the Depth-…☆20Aug 19, 2026Updated 2 weeks ago
- Transformers components but in Triton☆34May 9, 2025Updated last year
- A research workbench for developing and testing attacks against large language models, with a focus on prompt injection vulnerabilities a…☆60Jul 24, 2026Updated last month
- ☆33Nov 27, 2023Updated 2 years ago
- Physics of Language Models: Part 4.2, Canon Layers at Scale where Synthetic Pretraining Resonates in Reality☆360Aug 21, 2026Updated last week
- mHC-lite: You Don’t Need 20 Sinkhorn-Knopp Iterations☆94Jan 12, 2026Updated 7 months ago
- LCA-on-the-line (ICML 2024 Oral)☆14Feb 13, 2025Updated last year
- Official implementation of "Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought" (NeurIPS 2025)☆44Oct 8, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆56Dec 12, 2023Updated 2 years ago
- The official code of "Mano: Restriking Manifold Optimization for LLM Training".☆25Jun 1, 2026Updated 3 months ago
- The official repository of the paper "On the Exploitability of Instruction Tuning".☆71Feb 5, 2024Updated 2 years ago
- Building Tiny Recursive Models from Scratch☆15Oct 9, 2025Updated 10 months ago
- Algorithms for approximate attention in LLMs☆22Apr 14, 2025Updated last year
- Official repository for Parallax (Parameterized Local Linear Attention)☆68Jul 30, 2026Updated last month
- latent context language models☆77Updated this week