Large Language Model (LLM) Serving Paper and Resource List
☆30Aug 25, 2026Updated 3 weeks ago
Alternatives and similar repositories for Awesome-LLM-Serving
Users that are interested in Awesome-LLM-Serving are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Local-first, open-source Claude Science alternative. Web/ Desktop App. Claude Code/ Codex.☆21Jul 1, 2026Updated 2 months ago
- ☆19Jun 17, 2022Updated 4 years ago
- ☆20May 10, 2025Updated last year
- ☆63Updated this week
- ☆24Jun 1, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Research prototype of PRISM — a cost-efficient multi-LLM serving system with flexible time- and space-based GPU sharing.☆77Mar 17, 2026Updated 6 months ago
- ISCA-2025☆26Mar 3, 2026Updated 6 months ago
- Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation☆42Sep 7, 2026Updated last week
- 七夕孤寡助手☆13Aug 7, 2021Updated 5 years ago
- Oversubscription of GPU Memory through Transparent Swapping☆15Mar 27, 2015Updated 11 years ago
- Memory footprint reduction for transformer models☆11Jan 24, 2023Updated 3 years ago
- ☆13Sep 19, 2024Updated 2 years ago
- An HLS based winograd systolic CNN accelerator☆54Jul 18, 2021Updated 5 years ago
- ☆15Jun 26, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Fast and Flexible FPGA development using Hierarchical Partial Reconfiguration (FPT 2022)☆16Mar 21, 2024Updated 2 years ago
- ☆16May 27, 2026Updated 3 months ago
- HLS project modeling various sparse accelerators.☆12Jan 11, 2022Updated 4 years ago
- ☆14Nov 7, 2024Updated last year
- ☆24Apr 10, 2022Updated 4 years ago
- ☆74Feb 16, 2023Updated 3 years ago
- Python C++ Code Manager☆16Sep 29, 2024Updated last year
- a mllm inference engine for academic research☆22Jan 30, 2026Updated 7 months ago
- ☆13Aug 1, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- DUTH RISC V Microprocessor for High Level Synthesis☆10Jun 23, 2023Updated 3 years ago
- ☆10Mar 20, 2021Updated 5 years ago
- Herald: Accelerating Neural Recommendation Training with Embedding Scheduling (NSDI 2024)☆23May 9, 2024Updated 2 years ago
- CNN simd based accelerator using Vitis HLS☆12Jul 15, 2022Updated 4 years ago
- ☆13Mar 6, 2023Updated 3 years ago
- ☆10Jul 5, 2023Updated 3 years ago
- ☆15Jul 7, 2020Updated 6 years ago
- Artifact of OSDI '24 paper, ”Llumnix: Dynamic Scheduling for Large Language Model Serving“☆66Jun 5, 2024Updated 2 years ago
- An HBM FPGA based SpMV Accelerator☆19Aug 29, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆11Jan 25, 2023Updated 3 years ago
- 面向可信执行环境的OS。☆12May 9, 2025Updated last year
- Implementation of paper "GraphACT: Accelerating GCN Training on CPU-FPGA Heterogeneous Platform".☆12Jun 25, 2020Updated 6 years ago
- ☆12Mar 5, 2025Updated last year
- ☆13Jun 20, 2023Updated 3 years ago
- introduce AI infra knowledges. 人工智能系统基础架构知识库☆17Jun 4, 2023Updated 3 years ago
- Recent Advances on MLLM's Reasoning Ability☆26Apr 11, 2025Updated last year