通过实验对比LLM推理中Prefill和Decoding阶段的吞吐量差异,揭示性能瓶颈,解释PD分离优化技术的原理。包含CUDA和Apple MPS (M系列芯片) 的测试脚本。
☆23May 22, 2025Updated last year
Alternatives and similar repositories for LLM-Prefill-Decode-Benchmark
Users that are interested in LLM-Prefill-Decode-Benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- LaTeX Examples Document Source☆12Apr 9, 2024Updated 2 years ago
- ☆14Feb 2, 2021Updated 5 years ago
- A bert baseline for DocRED☆18Oct 12, 2022Updated 3 years ago
- A Model Agnostic function to directly remove specified layers from the LLM☆11May 23, 2024Updated 2 years ago
- ☆13Sep 8, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Official code for Guiding Language Model Math Reasoning with Planning Tokens☆21Feb 29, 2024Updated 2 years ago
- The official repo for "VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search" [EMNLP25]☆39Feb 1, 2026Updated 7 months ago
- This repository contains papers for a comprehensive survey on accelerated generation techniques in Large Language Models (LLMs).☆11May 24, 2024Updated 2 years ago
- ☆11Nov 19, 2024Updated last year
- AI Hedge Fund Repo integrate with DeepSeek V3 and R1 hosted on SiliconFlow.☆13Feb 3, 2025Updated last year
- a simple new ISA nnISA and nnSOC nnCPU nnAs nnCc☆10Mar 15, 2020Updated 6 years ago
- Mathematical expression evaluator with just in time code generation.☆12Apr 7, 2013Updated 13 years ago
- ROCm Driver RDMA Peer to Peer Support☆22Mar 21, 2019Updated 7 years ago
- Made a CPU in Logisim when I was 14 (2009), and wrote a naive assembler and compiler for it in Flash. The CPU's design is inspired by Don…☆10Sep 30, 2016Updated 9 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- An Empirical Study of Memorization in NLP (ACL 2022)☆13Jun 22, 2022Updated 4 years ago
- Distributed SDDMM Kernel☆13Jul 8, 2022Updated 4 years ago
- Optimised x86-64 gzip decompressor☆37Feb 26, 2018Updated 8 years ago
- edgetpu-native from https://coral.googlesource.com/edgetpu-native☆10Apr 15, 2019Updated 7 years ago
- "A relativist is an individual who doesn't know the difference between an adjective and an adverb." ― Bill Gaede☆23Dec 3, 2020Updated 5 years ago
- llvmのAZ Processor Backend☆11Oct 29, 2013Updated 12 years ago
- 8-bit RISC Processor on Logisim☆14Oct 1, 2020Updated 5 years ago
- Minimal FPGA Processor Core for Stack-based CPU for CPLDs Using Bit-Serial Architecture☆18Sep 6, 2013Updated 13 years ago
- CUDA C simple application for Nvidia's GPU☆11Jun 7, 2022Updated 4 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆10Feb 8, 2023Updated 3 years ago
- Convex Formulation of Multiple Instance Learning from Positive and Unlabeled Bags☆10Apr 28, 2018Updated 8 years ago
- Source code and data for the EDM 2022 paper☆12May 16, 2022Updated 4 years ago
- 北京理工大学大四小学期计算机组成原理部分☆12Sep 24, 2020Updated 5 years ago
- Some microbenchmarks and design docs before commencement☆11Feb 1, 2021Updated 5 years ago
- Human Resource Management App☆14Aug 5, 2016Updated 10 years ago
- Learn how to use Shiboken2 with your own custom Qt based library☆11Nov 3, 2021Updated 4 years ago
- Example of applying CUDA graphs to LLaMA-v2☆11Aug 25, 2023Updated 3 years ago
- ☆13Jan 7, 2025Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆15Mar 11, 2026Updated 6 months ago
- Implementation of the first neural natural logic paper on natural language inference☆10Oct 31, 2022Updated 3 years ago
- Tasks and tutorials using Graphore's IPU with Hugging Face. Originally at https://github.com/gradient-ai/Graphcore-HuggingFace☆16Mar 12, 2024Updated 2 years ago
- 🚀 LLM inference optimization simulator, modeling compute-bound prefill and memory-bound decode phases.☆13Jul 12, 2025Updated last year
- 华科七边形,欢迎各位朋友的指导与交流。☆32Nov 9, 2024Updated last year
- Customized Inference Engine for Multiverse Models☆26Jun 27, 2025Updated last year
- LLM Evaluation Framework for Hardware Design Using Python-Embedded DSLs☆18Aug 26, 2024Updated 2 years ago