Large Language Model (LLM) Serving Paper and Resource List
☆29Jul 16, 2026Updated 3 weeks ago
Alternatives and similar repositories for Awesome-LLM-Serving
Users that are interested in Awesome-LLM-Serving are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Local-first, open-source Claude Science alternative. Web/ Desktop App. Claude Code/ Codex.☆21Jul 1, 2026Updated last month
- ☆18Jun 17, 2022Updated 4 years ago
- ☆62Jul 1, 2025Updated last year
- ☆23Jun 1, 2025Updated last year
- [KDD 2025] The source code for UQABench☆12Aug 18, 2025Updated 11 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ISCA-2025☆25Mar 3, 2026Updated 5 months ago
- Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation☆42Jul 20, 2026Updated 2 weeks ago
- 七夕孤寡助手☆13Aug 7, 2021Updated 5 years ago
- Implementation of algorithms for memory optimized deep neural network training☆10Jul 23, 2020Updated 6 years ago
- Oversubscription of GPU Memory through Transparent Swapping☆15Mar 27, 2015Updated 11 years ago
- ☆13Sep 19, 2024Updated last year
- Correlated Low-rank Structure (CoLR) for Federated Recommendation System☆13May 31, 2026Updated 2 months ago
- An HLS based winograd systolic CNN accelerator☆54Jul 18, 2021Updated 5 years ago
- ☆15Jun 26, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Fast and Flexible FPGA development using Hierarchical Partial Reconfiguration (FPT 2022)☆15Mar 21, 2024Updated 2 years ago
- ☆16May 27, 2026Updated 2 months ago
- HLS project modeling various sparse accelerators.☆12Jan 11, 2022Updated 4 years ago
- ☆14Nov 7, 2024Updated last year
- ☆24Apr 10, 2022Updated 4 years ago
- ☆74Feb 16, 2023Updated 3 years ago
- Python C++ Code Manager☆15Sep 29, 2024Updated last year
- a mllm inference engine for academic research☆21Jan 30, 2026Updated 6 months ago
- ☆13Aug 1, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Official implementation for "K-Forcing: Joint Next-K-Token Decoding via Push-Forward Language Modeling"☆18Jun 14, 2026Updated last month
- DUTH RISC V Microprocessor for High Level Synthesis☆10Jun 23, 2023Updated 3 years ago
- ☆63Mar 24, 2025Updated last year
- CNN simd based accelerator using Vitis HLS☆11Jul 15, 2022Updated 4 years ago
- (ACL2025 oral) SCOPE: Optimizing KV Cache Compression in Long-context Generation☆36May 28, 2025Updated last year
- ☆13Mar 6, 2023Updated 3 years ago
- ☆10Jul 5, 2023Updated 3 years ago
- Implementation of vDNN++; an improvement over vDNN☆18Dec 7, 2018Updated 7 years ago
- ☆15Jul 7, 2020Updated 6 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- This is a cross-chip platform collection of operators and a unified neural network library.☆17Nov 3, 2023Updated 2 years ago
- Thinking is hard - automate it☆18Aug 24, 2022Updated 3 years ago
- An HBM FPGA based SpMV Accelerator☆19Aug 29, 2024Updated last year
- ☆11Jan 25, 2023Updated 3 years ago
- Official repository of paper [FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic, NeurIPS 2025]☆21Dec 2, 2025Updated 8 months ago
- [CVPRW 2026 Oral] Less Detail, Better Answers: Degradation-Driven Prompting for VQA☆20Apr 25, 2026Updated 3 months ago
- 面向可信执行环境的OS。☆12May 9, 2025Updated last year