Landing repository for the paper "Predicting the Order of Upcoming Tokens Improves Language Modeling"
☆48May 13, 2026Updated 3 months ago
Alternatives and similar repositories for token-order-prediction
Users that are interested in token-order-prediction are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17Jul 22, 2026Updated 3 weeks ago
- An MLX implementation of Meta AI's ESM-2 protein language model☆16Aug 16, 2025Updated last year
- Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types☆32Jul 16, 2025Updated last year
- ☆17Jul 31, 2025Updated last year
- The official repo for “Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem” [EMNLP25]☆33Sep 1, 2025Updated 11 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICML 2026] Esoteric Language Models☆124Jul 13, 2026Updated last month
- ☆21Jun 12, 2025Updated last year
- ☆23Jun 16, 2026Updated 2 months ago
- Official PyTorch implementation and models for paper "Diffusion Beats Autoregressive in Data-Constrained Settings". We find diffusion mod…☆128Jan 10, 2026Updated 7 months ago
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 4 months ago
- ☆55Sep 10, 2025Updated 11 months ago
- An official implementation of Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards☆36Oct 3, 2025Updated 10 months ago
- This is an implementation of the paper "Are We Done with Object-Centric Learning?"☆14Jun 21, 2026Updated last month
- Landing repository for the paper "Softpick: No Attention Sink, No Massive Activations with Rectified Softmax"☆92Sep 12, 2025Updated 11 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆14Jan 22, 2025Updated last year
- LLMBind: A Unified Modality-Task Integration Framework☆19Jun 16, 2024Updated 2 years ago
- The official repository for SkyLadder: Better and Faster Pretraining via Context Window Scheduling☆43Dec 29, 2025Updated 7 months ago
- This repository is associated with the research paper titled ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large…☆15Jun 4, 2025Updated last year
- Efficient non-uniform quantization with GPTQ for GGUF☆64Sep 17, 2025Updated 11 months ago
- ☆14Updated this week
- Code and datasets for "Text encoders are performance bottlenecks in contrastive vision-language models". Coming soon!☆11May 24, 2023Updated 3 years ago
- [COLM 2025] Official code for "When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoni…☆15Oct 31, 2025Updated 9 months ago
- Code for "DynaGuard: A Dynamic Guardrail Model With User-Defined Policies."☆23Nov 3, 2025Updated 9 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Code for boomerang distillation enables zero-shot model size interpolation.☆22Jul 10, 2026Updated last month
- Subliminal learning in LLMs: language models can transmit hidden preferences through seemingly unrelated training data.☆25Nov 9, 2025Updated 9 months ago
- Code for Generalized Entropy Regularization paper☆14May 2, 2020Updated 6 years ago
- [BabyLM@EMNLP 2025 - Challenge Award] Official Implementation: Masked Diffusion Language Models with Frequency-Informed Training☆16Dec 17, 2025Updated 8 months ago
- [NeurIPS '25] Multi-Token Prediction Needs Registers☆32Dec 14, 2025Updated 8 months ago
- Storing long contexts in tiny caches with self-study☆314Mar 23, 2026Updated 4 months ago
- Extending the Context of Pretrained LLMs by Dropping Their Positional Embedding☆220Jan 12, 2026Updated 7 months ago
- [ICLR 2026] Adapting Self-Supervised Representations as a Latent Space for Efficient Generation☆60Apr 24, 2026Updated 3 months ago
- MLX Implementation of Recursive Reasoning with Tiny Networks☆79Oct 11, 2025Updated 10 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Official implementation of "Token Perturbation Guidance for Diffusion Models" [NeurIPS 2025]☆17May 19, 2026Updated 2 months ago
- An automated data pipeline scaling RL to pretraining levels☆76Jun 2, 2026Updated 2 months ago
- Code for the paper "AsFT: Anchoring Safety During LLM Fune-Tuning Within Narrow Safety Basin".☆37Jul 10, 2025Updated last year
- The raw UserRL repo under construction☆115Jun 2, 2026Updated 2 months ago
- ☆22Jun 5, 2025Updated last year
- Reproducible and flexible LLM evaluations for scientific reasoning.☆30Jul 23, 2025Updated last year
- [ACL'26 Findings] Steering LLM Thinking with Budget Guidance☆34Feb 19, 2026Updated 5 months ago