Efficient LLM Inference Acceleration using Prompting
☆50Oct 22, 2024Updated last year
Alternatives and similar repositories for parallel-prompt-decoding
Users that are interested in parallel-prompt-decoding are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Codebase for the Progressive Mixed-Precision Decoding paper.☆22Jul 15, 2025Updated last year
- An innovative method expediting LLMs via streamlined semi-autoregressive generation and draft verification.☆28Apr 15, 2025Updated last year
- Official implementation of LittleBit (NeurIPS 2025) and its follow-up LittleBit-2 (ICML 2026)☆27May 6, 2026Updated 2 months ago
- List Flower resources☆12Feb 4, 2022Updated 4 years ago
- This repository contains papers for a comprehensive survey on accelerated generation techniques in Large Language Models (LLMs).☆11May 24, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Federated Learning for Artificial Intelligence and Machine Learning☆14Sep 11, 2025Updated 10 months ago
- [ICCAD 2025] Squant☆15Jul 3, 2025Updated last year
- A Hackable Quantization Library for PyTorch☆22Mar 29, 2021Updated 5 years ago
- ☆16Dec 9, 2023Updated 2 years ago
- ☆35Feb 10, 2025Updated last year
- BESA is a differentiable weight pruning technique for large language models.☆17Mar 4, 2024Updated 2 years ago
- ☆19Sep 29, 2024Updated last year
- The official implementation of TinyTrain [ICML '24]☆27Jul 19, 2024Updated 2 years ago
- ☆15Apr 11, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICML 2024] When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models☆35Jun 12, 2024Updated 2 years ago
- ☆22Apr 17, 2025Updated last year
- ☆11Feb 5, 2026Updated 5 months ago
- ☆93Aug 18, 2024Updated last year
- ☆15Apr 6, 2026Updated 3 months ago
- Fast and Robust Early-Exiting Framework for Autoregressive Language Models with Synchronized Parallel Decoding (EMNLP 2023 Long)☆67Sep 28, 2024Updated last year
- FaceGrabber is introduced in the following paper: D. Merget, T. Eckl, M. Schwörer, P. Tiefenbacher, and G. Rigoll, “Capturing Facial Vide…☆11Sep 7, 2016Updated 9 years ago
- Test scripts for exploring PyTorch JIT and quantization capability☆11Mar 8, 2021Updated 5 years ago
- Official evaluation models and configuration for Stage 1 of the SAIR Mathematics Distillation Challenge: Equational Theories.☆18Apr 19, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆10Aug 2, 2021Updated 4 years ago
- Code for "LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding", ACL 2024☆372Updated this week
- [ICML 2024 Oral] Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs☆130Jul 4, 2025Updated last year
- ☆19Mar 25, 2025Updated last year
- Tight Mutual Information Estimation With Contrastive Fenchel-Legendre Optimization☆11Nov 29, 2022Updated 3 years ago
- This repository contains the official implementation for the ECCV'22 paper, "SPIN: An Empirical Evaluation on Sharing Parameters of Isotr…☆20Sep 9, 2023Updated 2 years ago
- Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)☆401Apr 22, 2025Updated last year
- μNAS is a neural architecture search (NAS) system that designs small-yet-powerful microcontroller-compatible neural networks.☆84Jan 26, 2021Updated 5 years ago
- The official repo of continuous speculative decoding☆36Mar 28, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆56Jul 7, 2025Updated last year
- ☆26May 30, 2023Updated 3 years ago
- snn implementation for spike-deeplab and spike-fcn☆12Dec 31, 2022Updated 3 years ago
- ☆35Nov 18, 2025Updated 8 months ago
- Instruct-tuning LLaMA on consumer hardware with machine-translated data☆19Apr 17, 2023Updated 3 years ago
- Official code and data repository of MathChat: MathChat: Benchmarking Mathematical Reasoning and Instruction Following in Multi-Turn Inte…☆22Jun 3, 2024Updated 2 years ago
- Official codes for Scalable Infomin Learning, NeurIPS 2022☆14Feb 28, 2023Updated 3 years ago