[ICLR2025 Spotlightπ₯] Official Implementation of TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters
β595Feb 11, 2025Updated last year
Alternatives and similar repositories for TokenFormer
Users that are interested in TokenFormer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Minimal implementation of TokenFormer for inference and learningβ13Nov 6, 2024Updated last year
- [ECCV2024 Oralπ₯] Official Implementation of "GiT: Towards Generalist Vision Transformer through Universal Language Interface"β364Jan 14, 2025Updated last year
- Memory layers use a trainable key-value lookup mechanism to add extra parameters to a model without increasing FLOPs. Conceptually, sparsβ¦β379Dec 12, 2024Updated last year
- Code for BLT research paperβ2,059Nov 3, 2025Updated 9 months ago
- code for "Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion"β1,289Jul 6, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Quick implementation of nGPT, learning entirely on the hypersphere, from NvidiaAIβ300Jun 3, 2025Updated last year
- Next-Token Prediction is All You Needβ2,442Jan 12, 2026Updated 7 months ago
- Helpful tools and examples for working with flex-attentionβ1,233Updated this week
- π Efficient implementations for emerging model architecturesβ5,682Updated this week
- Layer-Condensed KV cache w/ 10 times larger batch size, fewer params and less computation. Dramatic speed up with better task performanceβ¦β156Apr 7, 2025Updated last year
- FlexAttention w/ FlashAttention3 Supportβ27Oct 5, 2024Updated last year
- [ICLR 2025 & COLM 2025] Official PyTorch implementation of the Forgetting Transformer and Adaptive Computation Pruningβ152Feb 25, 2026Updated 6 months ago
- [ICLR'25 Oral] Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Thinkβ1,706Mar 16, 2025Updated last year
- A suite of image and video neural tokenizersβ1,731Feb 11, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- β53Jun 24, 2025Updated last year
- Code for Adam-mini: Use Fewer Learning Rates To Gain More https://arxiv.org/abs/2406.16793β459May 13, 2025Updated last year
- β93Aug 18, 2024Updated 2 years ago
- Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Modelsβ343Feb 23, 2025Updated last year
- Mamba SSM architectureβ18,796Jul 22, 2026Updated last month
- Block Transformer: Global-to-Local Language Modeling for Fast Inference (NeurIPS 2024)β167Apr 13, 2025Updated last year
- PyTorch implementation of MAR+DiffLoss https://arxiv.org/abs/2406.11838β1,952Feb 20, 2026Updated 6 months ago
- Code for NeurIPS 2024 Spotlight: "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations"β93Oct 30, 2024Updated last year
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projectionβ1,701Oct 28, 2024Updated last year
- Open source password manager - Proton Pass β’ AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- β285Jun 6, 2025Updated last year
- Pretraining and inference code for a large-scale depth-recurrent language modelβ920Dec 29, 2025Updated 8 months ago
- Official PyTorch Implementation of "Scalable Diffusion Models with Transformers"β8,689May 31, 2024Updated 2 years ago
- Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clustersβ135Dec 3, 2024Updated last year
- Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.β2,103Jul 29, 2024Updated 2 years ago
- Reference implementation of "Softmax Attention with Constant Cost per Token" (Heinsen, 2024)β25Jun 6, 2024Updated 2 years ago
- Code for exploring Based models from "Simple linear attention language models balance the recall-throughput tradeoff"β258Jun 6, 2025Updated last year
- β19Dec 4, 2025Updated 8 months ago
- Annotated version of the Mamba paperβ502Feb 27, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Efficient Triton Kernels for LLM Trainingβ6,597Updated this week
- Official PyTorch Implementation of "SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers"β1,206Dec 22, 2025Updated 8 months ago
- Repo for "Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture"β563Dec 28, 2024Updated last year
- Muon is Scalable for LLM Trainingβ1,541Aug 3, 2025Updated last year
- Official implementation of Next Block Prediction: Video Generation via Semi-Autoregressive Modelingβ42Feb 12, 2025Updated last year
- β311Apr 23, 2025Updated last year
- Autoregressive Model Beats Diffusion: π¦ Llama for Scalable Image Generationβ1,966Aug 15, 2024Updated 2 years ago