Landing repository for the paper "Predicting the Order of Upcoming Tokens Improves Language Modeling"
☆48May 13, 2026Updated 4 months ago
Alternatives and similar repositories for token-order-prediction
Users that are interested in token-order-prediction are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17Updated this week
- An MLX implementation of Meta AI's ESM-2 protein language model☆16Aug 16, 2025Updated last year
- Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types☆32Jul 16, 2025Updated last year
- ☆17Sep 11, 2026Updated 2 weeks ago
- The official repo for “Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem” [EMNLP25]☆33Sep 1, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICML 2026] Esoteric Language Models☆125Jul 13, 2026Updated 2 months ago
- ☆23Jun 12, 2025Updated last year
- ☆24Jun 16, 2026Updated 3 months ago
- Official PyTorch implementation and models for paper "Diffusion Beats Autoregressive in Data-Constrained Settings". We find diffusion mod…☆128Jan 10, 2026Updated 8 months ago
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 5 months ago
- An official implementation of Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards☆36Oct 3, 2025Updated 11 months ago
- ☆57Sep 10, 2025Updated last year
- ☆14Jan 22, 2025Updated last year
- LLMBind: A Unified Modality-Task Integration Framework☆19Jun 16, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- The official repository for SkyLadder: Better and Faster Pretraining via Context Window Scheduling☆43Dec 29, 2025Updated 8 months ago
- This repository is associated with the research paper titled ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large…☆15Jun 4, 2025Updated last year
- Efficient non-uniform quantization with GPTQ for GGUF☆66Sep 17, 2025Updated last year
- ☆14Aug 14, 2026Updated last month
- Code and datasets for "Text encoders are performance bottlenecks in contrastive vision-language models". Coming soon!☆11May 24, 2023Updated 3 years ago
- Code for boomerang distillation enables zero-shot model size interpolation.☆23Jul 10, 2026Updated 2 months ago
- Code for Generalized Entropy Regularization paper☆14May 2, 2020Updated 6 years ago
- [BabyLM@EMNLP 2025 - Challenge Award] Official Implementation: Masked Diffusion Language Models with Frequency-Informed Training☆16Dec 17, 2025Updated 9 months ago
- ☆15Oct 4, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [NeurIPS '25] Multi-Token Prediction Needs Registers☆32Dec 14, 2025Updated 9 months ago
- Implementation of Bitune: Bidirectional Instruction-Tuning☆27Jun 19, 2025Updated last year
- Extending the Context of Pretrained LLMs by Dropping Their Positional Embedding☆224Jan 12, 2026Updated 8 months ago
- [ICLR 2026] Adapting Self-Supervised Representations as a Latent Space for Efficient Generation☆62Apr 24, 2026Updated 5 months ago
- MLX Implementation of Recursive Reasoning with Tiny Networks☆79Oct 11, 2025Updated 11 months ago
- An automated data pipeline scaling RL to pretraining levels☆76Jun 2, 2026Updated 3 months ago
- Distributed file system☆13May 10, 2011Updated 15 years ago
- Code for the paper "Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns"☆19Mar 15, 2024Updated 2 years ago
- implementation of https://arxiv.org/pdf/2312.09299☆21Jul 3, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code for the paper "AsFT: Anchoring Safety During LLM Fune-Tuning Within Narrow Safety Basin".☆38Jul 10, 2025Updated last year
- ☆22Jun 5, 2025Updated last year
- Reproducible and flexible LLM evaluations for scientific reasoning.☆32Jul 23, 2025Updated last year
- The raw UserRL repo under construction☆120Jun 2, 2026Updated 3 months ago
- [ACL'26 Findings] Steering LLM Thinking with Budget Guidance☆34Feb 19, 2026Updated 7 months ago
- official code repo for paper "Merging Models on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model Merging"☆25Oct 11, 2025Updated 11 months ago
- [ICLR 26] The official code repository for the paper "Mirage or Method? How Model–Task Alignment Induces Divergent RL Conclusions".☆19Feb 9, 2026Updated 7 months ago