通过实验对比LLM推理中Prefill和Decoding阶段的吞吐量差异,揭示性能瓶颈,解释PD分离优化技术的原理。包含CUDA和Apple MPS (M系列芯片) 的测试脚本。
☆23May 22, 2025Updated last year
Alternatives and similar repositories for LLM-Prefill-Decode-Benchmark
Users that are interested in LLM-Prefill-Decode-Benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ROCm Command Line Profiler - Updated moved to https://github.com/GPUOpen-Tools/RCP☆10Aug 24, 2017Updated 8 years ago
- LaTeX Examples Document Source☆11Apr 9, 2024Updated 2 years ago
- ☆13Sep 8, 2024Updated last year
- ToyLLM: Learning LLM from Scratch☆25Updated this week
- In-depth tutorials and examples on LLM training and inference infrastructure, such as, Pytorch, Fairscale, Nvidia AI Modules (cuDNN, tens…☆22May 19, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- AI Hedge Fund Repo integrate with DeepSeek V3 and R1 hosted on SiliconFlow.☆12Feb 3, 2025Updated last year
- ☆11Nov 19, 2024Updated last year
- ☆14Feb 12, 2024Updated 2 years ago
- QCD for Intel Xeon Phi and Xeon processors☆14Updated this week
- a simple new ISA nnISA and nnSOC nnCPU nnAs nnCc☆10Mar 15, 2020Updated 6 years ago
- ROCm Driver RDMA Peer to Peer Support☆22Mar 21, 2019Updated 7 years ago
- 👩💻 Code for the ACL paper "Detecting Edit Failures in LLMs: An Improved Specificity Benchmark"☆20Jan 19, 2024Updated 2 years ago
- 基于FPGA-Pynq的车牌识别系统。The LPR system of FPGA-Pynq☆13Mar 22, 2019Updated 7 years ago
- demonstration for our ACL 2018 paper, "On the Practical Computational Power of Finite Precision RNNs for Language Recognition"☆11May 26, 2019Updated 7 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Code for "Interpreting Word Embeddings with Eigenvector Analysis" https://openreview.net/forum?id=rJfJiR5ooX.☆16Oct 16, 2019Updated 6 years ago
- Distributed SDDMM Kernel☆13Jul 8, 2022Updated 4 years ago
- Optimised x86-64 gzip decompressor☆33Feb 26, 2018Updated 8 years ago
- ☆11Dec 31, 2020Updated 5 years ago
- llvmのAZ Processor Backend☆11Oct 29, 2013Updated 12 years ago
- [ICLR 2024] Unveiling the Pitfalls of Knowledge Editing for Large Language Models☆22Jun 13, 2024Updated 2 years ago
- 8-bit RISC Processor on Logisim☆14Oct 1, 2020Updated 5 years ago
- ☆21Apr 18, 2024Updated 2 years ago
- Multi-Level Adversarial for Cross-lingual Name Tagging☆12Jun 18, 2020Updated 6 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- CUDA C simple application for Nvidia's GPU☆11Jun 7, 2022Updated 4 years ago
- Large-scale Exploration of Neural Relation Classification Architectures☆11Nov 15, 2018Updated 7 years ago
- Source code and data for the EDM 2022 paper☆12May 16, 2022Updated 4 years ago
- 北京理工大学大四小学期计算机组成原理部分☆12Sep 24, 2020Updated 5 years ago
- ☆16Aug 19, 2024Updated last year
- The training framework of PCMind-2.1-Kaiyuan-2B built on MindFormers☆15Dec 9, 2025Updated 7 months ago
- Personal Learning Notes☆25Jan 17, 2019Updated 7 years ago
- ☆13Jan 7, 2025Updated last year
- Example of applying CUDA graphs to LLaMA-v2☆11Aug 25, 2023Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆15Mar 11, 2026Updated 4 months ago
- An Efficient and Versatile Inference Engine for Distributed LLM Serving☆66Updated this week
- ☆45Oct 11, 2025Updated 9 months ago
- Tasks and tutorials using Graphore's IPU with Hugging Face. Originally at https://github.com/gradient-ai/Graphcore-HuggingFace☆16Mar 12, 2024Updated 2 years ago
- [AAAI 2025] Assessing the Creativity of LLMs in Proposing Novel Solutions to Mathematical Problems☆13May 5, 2025Updated last year
- 华科七边形,欢迎各位朋友的指导与交流。☆32Nov 9, 2024Updated last year
- Customized Inference Engine for Multiverse Models☆26Jun 27, 2025Updated last year