This repository is the official implementation of "Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE" [ACL 2026 Main Accepted]
☆37Oct 5, 2025Updated 9 months ago
Alternatives and similar repositories for Jakiro
Users that are interested in Jakiro are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆25Mar 15, 2023Updated 3 years ago
- ☆15May 27, 2025Updated last year
- Anatomy of High-Performance GEMM with Online Fault Tolerance on GPUs☆14Apr 3, 2025Updated last year
- 📰 Must-read papers and blogs on Speculative Decoding ⚡️☆1,283Jun 27, 2026Updated last month
- [ACL 2026 (Main)] LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification☆84Jul 14, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Pixels, Patterns, but no Poetry: To See the World like Humans☆18Aug 11, 2025Updated 11 months ago
- ☆20Jun 17, 2024Updated 2 years ago
- GCAPS: GPU Context-Aware Preemptive Scheduling Approach☆16Mar 22, 2026Updated 4 months ago
- BigBang-Proton is a LLM pretrained on cross-scale, cross-structure, cross-discipline real-world scientific tasks to construct a scienti…☆21Nov 8, 2025Updated 8 months ago
- [ICML 2024] When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models☆35Jun 12, 2024Updated 2 years ago
- [ICML‘25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training an…☆13Apr 17, 2025Updated last year
- ☆24Mar 7, 2025Updated last year
- Source code of paper ''KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing''☆31Oct 24, 2024Updated last year
- GPU topology-aware scheduler☆13Jul 7, 2017Updated 9 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Landing page + leaderboard for SWE-Bench benchmark☆15Mar 29, 2026Updated 4 months ago
- This project leverages advanced AI agents from crewAI to assist doctors in diagnosing medical conditions and recommending treatment plans…☆15Nov 16, 2024Updated last year
- ☆47May 27, 2025Updated last year
- ☆11Oct 29, 2022Updated 3 years ago
- Official resources of "The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reaso…☆20Jun 13, 2025Updated last year
- [COLM 2024] TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding☆281Aug 31, 2024Updated last year
- Official Implementation of LANTERN (ICLR'25) and LANTERN++(ICLRW-SCOPE'25)☆21Mar 5, 2025Updated last year
- ☆18Jul 31, 2025Updated 11 months ago
- Official PyTorch implementation of [PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation](https://arxiv.org/abs…☆25Jan 25, 2026Updated 6 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Streamlit template for building SMART on FHIR apps in the Cerner ecosystem.☆11Sep 22, 2023Updated 2 years ago
- ITKGrowCut is a remote module for ITK. It segments a 3D image from user-provided foreground and background seeds.☆16May 27, 2026Updated 2 months ago
- Fast inference from large lauguage models via speculative decoding☆921Aug 22, 2024Updated last year
- Simple MPR medical imaging viewer using VTK 6.0.1 and Qt 5.2.1☆10Aug 15, 2014Updated 11 years ago
- Simple implementation of Speculative Sampling in NumPy for GPT-2.☆99Aug 20, 2023Updated 2 years ago
- A practical guide to high-performance gluon kernel development on AMD GFX9 GPUs.☆41Updated this week
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).☆2,481Feb 20, 2026Updated 5 months ago
- ☆23Oct 10, 2025Updated 9 months ago
- ☆40Nov 18, 2025Updated 8 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Automating analysis from trace files☆85Updated this week
- 3x Faster Inference; Unofficial implementation of EAGLE Speculative Decoding☆85Jul 3, 2025Updated last year
- Ongoing research training transformer models at scale☆43Updated this week
- [ICLR2025] Code and data for paper: Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasonin…☆45Mar 10, 2025Updated last year
- 本科毕设:基于VTK的三维可视化平台☆12Jun 13, 2019Updated 7 years ago
- Simulating Distributed Training at Scale☆14Sep 15, 2025Updated 10 months ago
- Can LLMs Convert Graphs to Text-Attributed Graphs? NAACL 25☆17Mar 7, 2025Updated last year