[ICML 2025] |TokenSwift: Lossless Acceleration of Ultra Long Sequence Generation
☆126May 19, 2025Updated last year
Alternatives and similar repositories for TokenSwift
Users that are interested in TokenSwift are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Repo for ReflectEvo☆21Jun 16, 2025Updated last year
- [NeurIPS 2024] | An Efficient Recipe for Long Context Extension via Middle-Focused Positional Encoding☆22Oct 10, 2024Updated last year
- TMLR | This survey presents a comprehensive and structured synthesis of memory in LLMs and MLLMs, organizing the literature into a cohesi…☆37Jan 28, 2026Updated 5 months ago
- [ICLR 2026] RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling☆39Feb 25, 2026Updated 5 months ago
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆23Jul 14, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆15Apr 14, 2026Updated 3 months ago
- [ICML 2026] Reasoning in Parallelism via Self-Distilled RL☆113Jun 28, 2026Updated 3 weeks ago
- [EMNLP 2024] A Video Chat Agent with Temporal Prior☆33Mar 2, 2025Updated last year
- Evals Harness for $OneMillion-Bench☆48Apr 21, 2026Updated 3 months ago
- Language Modeling Research Hub, a comprehensive compendium for enthusiasts and scholars delving into the fascinating realm of language mo…☆19Mar 19, 2025Updated last year
- Official Repo of LangSuitE☆85Aug 15, 2024Updated last year
- Official Repository of LatentSeek☆85Jun 6, 2025Updated last year
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆18Apr 2, 2025Updated last year
- [ICCV 2025] Official Repository of VideoLLaMB: Long Video Understanding with Recurrent Memory Bridges☆87Feb 27, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ACL 2024 | LooGLE: Long Context Evaluation for Long-Context Language Models☆199Oct 8, 2024Updated last year
- 【ICLR 2025 🔥】The code for Consistent In-Context Editing, an approach for tuning language models through contextual distributions, overco…☆56Apr 2, 2025Updated last year
- ☆30May 24, 2025Updated last year
- A high-fidelity, general-purpose platform for embodied agent training and testing.☆189Jan 12, 2026Updated 6 months ago
- Implementation of "Decoding-time Realignment of Language Models", ICML 2024.☆21Jun 17, 2024Updated 2 years ago
- [NeurIPS 2025] Simple extension on vLLM to help you speed up reasoning model without training.☆232May 31, 2025Updated last year
- ☆17Apr 9, 2025Updated last year
- FastCuRL: Curriculum Reinforcement Learning with Stage-wise Context Scaling for Efficient LLM Reasoning (EMNLP 2025)☆61Oct 10, 2025Updated 9 months ago
- Recent advancements propelled by large language models (LLMs), encompassing an array of domains including Vision, Audio, Agent, Robotics,…☆128May 7, 2026Updated 2 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- This repository includes code and materials for the paper "Efficient PRM Training Data Synthesis via Formal Verification" (ACL 2026 Findi…☆19Apr 7, 2026Updated 3 months ago
- Layer-Condensed KV cache w/ 10 times larger batch size, fewer params and less computation. Dramatic speed up with better task performance…☆157Apr 7, 2025Updated last year
- ☆47Jun 11, 2025Updated last year
- ☆20Jun 17, 2024Updated 2 years ago
- [NeurIPS 2024] Fast Best-of-N Decoding via Speculative Rejection☆56Oct 29, 2024Updated last year
- [ACL 2025 Findings] Implicit Reasoning in Transformers is Reasoning through Shortcuts☆18Mar 11, 2025Updated last year
- ☆65Mar 30, 2026Updated 3 months ago
- GEAR: An Efficient KV Cache Compression Recipefor Near-Lossless Generative Inference of LLM☆183Jul 12, 2024Updated 2 years ago
- Official Implementation of SAM-Decoding: Speculative Decoding via Suffix Automaton☆52May 12, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- a fully open-source implementation of a GPT-4o-like speech-to-speech video understanding model.☆38Apr 7, 2025Updated last year
- Work in progress.☆80Nov 25, 2025Updated 8 months ago
- ☆24Mar 7, 2025Updated last year
- Official repository of paper "Context-DPO: Aligning Language Models for Context-Faithfulness"☆23Feb 17, 2025Updated last year
- [ACL 2023] VSTAR is a multimodal dialogue dataset with scene and topic transition information☆16Oct 27, 2024Updated last year
- The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism☆31Jul 17, 2024Updated 2 years ago
- ☆36Jun 5, 2025Updated last year