Compile programs directly into transformer weights. Includes a 2D convex-hull KV cache with O(log n) inference.
☆215Jun 1, 2026Updated 2 months ago
Alternatives and similar repositories for transformer-vm
Users that are interested in transformer-vm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Method for Long Context RLMs using verifiable Lambda Calculus☆305Apr 24, 2026Updated 3 months ago
- My reasearch of losslessly compressing LLM weights.☆63Jul 23, 2026Updated last month
- AI-Driven Scientific, Algorithmic, and Systems Discovery☆616Updated this week
- ☆108Aug 11, 2026Updated last week
- SR²AM: Efficient Agentic Reasoning Through Self-Regulated Simulative Planning☆21May 22, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Minimal and scalable research codebase in JAX, designed for rapid iteration on frontier research in LLM and other autoregressive models.☆562Updated this week
- Stable Looped Models and their Scaling Laws☆174May 17, 2026Updated 3 months ago
- Official implementation of the transformer (TF) architecture suggested in a paper entitled "Looped Transformers as Programmable Computers…☆43Apr 8, 2023Updated 3 years ago
- ☆17Feb 13, 2026Updated 6 months ago
- RND1: Scaling Diffusion Language Models☆187Aug 17, 2026Updated last week
- Automatically extract executable programs from pruned mechanistic circuits, extending OpenAI's Sparse Circuits☆73Nov 23, 2025Updated 9 months ago
- Various ML tidbits in Python/PyTorch and C++☆88Jun 8, 2026Updated 2 months ago
- 58 implementations of synthetic learning problems from Jürgen Schmidhuber's papers (1989-2025). Pure numpy, laptop-runnable, paper-compar…☆207May 25, 2026Updated 2 months ago
- An interaction combinator runtime☆18Sep 23, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Benchmarking long-horizon chain-of-thought reasoning.☆43Apr 20, 2026Updated 4 months ago
- Latent Program Network (from the "Searching Latent Program Spaces" paper)☆114Nov 25, 2025Updated 8 months ago
- 🤖 Complete reproduction of 'AlphaGo Moment for Model Architecture Discovery' using MLX-LM instead of GPT-4. Autonomous neural architectu…☆30Jul 27, 2025Updated last year
- Official implementation of Categorical Flow Maps on text.☆68Feb 16, 2026Updated 6 months ago
- A transformer that executes a one-instruction Turing-complete computer — two approaches: hand-coded weights (no training) and learned fro…☆41Mar 3, 2026Updated 5 months ago
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆38Jun 12, 2026Updated 2 months ago
- ☆802Apr 16, 2026Updated 4 months ago
- An unbounded n-gram language model on Tiny Shakespeare☆22Jan 21, 2026Updated 7 months ago
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 4 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Jax Codebase for Evolutionary Strategies at the Hyperscale☆369Feb 27, 2026Updated 5 months ago
- TEVO: evolve LM motifs cheaply, then validate them in downstream train.py loops.☆19Apr 18, 2026Updated 4 months ago
- Official JAX implementation of End-to-End Test-Time Training for Long Context☆677Feb 15, 2026Updated 6 months ago
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆375Jul 9, 2026Updated last month
- ☆1,257Apr 5, 2026Updated 4 months ago
- [ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.☆39Mar 5, 2026Updated 5 months ago
- Minimal open-source implementation of AlphaProof and HyperTree Proof Search.☆87May 13, 2026Updated 3 months ago
- ☆309Mar 14, 2026Updated 5 months ago
- Evolution Pretraining Fully in Int Formats☆182Feb 25, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Generalized Optimal Transport Attention with Trainable Priors☆72Updated this week
- ROSA+: RWKV's ROSA implementation with fallback statistical predictor☆36Oct 13, 2025Updated 10 months ago
- Official Project Page for HLA: Higher-order Linear Attention (https://arxiv.org/abs/2510.27258)☆103Jun 15, 2026Updated 2 months ago
- AdaSplash: Adaptive Sparse Flash Attention (aka Flash Entmax Attention)☆47May 20, 2026Updated 3 months ago
- Official implementation of DiscoGen, for "Procedural Generation of Algorithm Discovery Tasks in Machine Learning"☆50Aug 4, 2026Updated 2 weeks ago
- ☆73Mar 16, 2026Updated 5 months ago
- ☆148Jun 18, 2026Updated 2 months ago