This repository is the official implementation of "Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE" [ACL 2026 Main Accepted]
☆37Oct 5, 2025Updated 10 months ago
Alternatives and similar repositories for Jakiro
Users that are interested in Jakiro are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The four major frameworks for 3D point cloud sparse acceleration, which are currently mainstream, are compared. These include MIT-HAN-LAB…☆28Feb 8, 2025Updated last year
- This repository is the official implementation of "Partial Channel Network: Compute Fewer, Perform Better". [AAAI 2026 Accepted]☆41Feb 11, 2025Updated last year
- Official Implementation of "Learning Harmonized Representations for Speculative Sampling" (HASS)☆56Mar 14, 2025Updated last year
- ☆25Mar 15, 2023Updated 3 years ago
- ☆15May 27, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [NeurIPS 2025🔥:] EVODiff is an inference-time refinement method for diffusion models that improves sampling efficiency and generative f…☆32Feb 2, 2026Updated 6 months ago
- 📰 Must-read papers and blogs on Speculative Decoding ⚡️☆1,292Jun 27, 2026Updated last month
- ☆11Feb 3, 2025Updated last year
- [ACL 2026 (Main)] LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification☆86Jul 14, 2025Updated last year
- Pixels, Patterns, but no Poetry: To See the World like Humans☆18Aug 11, 2025Updated last year
- ☆20Jun 17, 2024Updated 2 years ago
- BigBang-Proton is a LLM pretrained on cross-scale, cross-structure, cross-discipline real-world scientific tasks to construct a scienti…☆21Nov 8, 2025Updated 9 months ago
- (ACL2025 oral) SCOPE: Optimizing KV Cache Compression in Long-context Generation☆36May 28, 2025Updated last year
- [ICLR 2024] This is the official PyTorch implementation of "QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Mod…☆31Mar 12, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICML‘25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training an…☆13Apr 17, 2025Updated last year
- [NeurIPS 2024] The official implementation of "Kangaroo: Lossless Self-Speculative Decoding for Accelerating LLMs via Double Early Exitin…☆73Jun 26, 2024Updated 2 years ago
- Awesome system papers for AI☆21Updated this week
- ☆24Mar 7, 2025Updated last year
- Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)☆404Apr 22, 2025Updated last year
- Source code of paper ''KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing''☆31Oct 24, 2024Updated last year
- Slowdown prediction module of Echo: Simulating Distributed Training at Scale☆13Jul 11, 2026Updated last month
- An ITK implementation of the GraphCut framework. See 'Graph cuts and efficient ND image segmentation' by Boykov and Funka-Lea and 'Intera…☆12Sep 18, 2017Updated 8 years ago
- Rhetorical sentence classification using LLMs☆11Oct 26, 2025Updated 9 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- This project leverages advanced AI agents from crewAI to assist doctors in diagnosing medical conditions and recommending treatment plans…☆15Nov 16, 2024Updated last year
- RL_Dynamic_Network_Reconfiguration☆10Apr 13, 2023Updated 3 years ago
- ☆47May 27, 2025Updated last year
- This repository offers models and an evaluation framework for short-term load forecasting, with a focus on transfer learning, deep learni…☆17Aug 5, 2026Updated last week
- ☆17Jul 31, 2025Updated last year
- [COLM 2024] TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding☆280Aug 31, 2024Updated last year
- Official Implementation of LANTERN (ICLR'25) and LANTERN++(ICLRW-SCOPE'25)☆22Mar 5, 2025Updated last year
- Official PyTorch implementation of [PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation](https://arxiv.org/abs…☆25Jan 25, 2026Updated 6 months ago
- ☆14Jan 14, 2020Updated 6 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ITKGrowCut is a remote module for ITK. It segments a 3D image from user-provided foreground and background seeds.☆16May 27, 2026Updated 2 months ago
- Fast inference from large lauguage models via speculative decoding☆923Aug 22, 2024Updated last year
- Simple MPR medical imaging viewer using VTK 6.0.1 and Qt 5.2.1☆10Aug 15, 2014Updated 12 years ago
- Simple implementation of Speculative Sampling in NumPy for GPT-2.☆99Aug 20, 2023Updated 2 years ago
- Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).☆2,502Feb 20, 2026Updated 5 months ago
- ☆40Nov 18, 2025Updated 9 months ago
- Historical shortest-path distance querying index by pruned landmark labeling☆10May 24, 2014Updated 12 years ago