Must-read papers on improving efficiency for LLM serving clusters
☆34May 28, 2025Updated last year
Alternatives and similar repositories for Awesome-Large-Scale-LLM-Serving
Users that are interested in Awesome-Large-Scale-LLM-Serving are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This is a displayer for Live2Dv3 model☆19Mar 11, 2020Updated 6 years ago
- ☆39Mar 17, 2025Updated last year
- A curated reading list for machine learning reliability research and practice☆32Sep 18, 2025Updated last year
- Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction | A tiny BERT model can tell you the verbosity of an …☆53Jun 1, 2024Updated 2 years ago
- Calibrating LLM Confidence by Probing Perturbed Representation Stability☆19Jul 5, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ICSE2021 Submission☆13Aug 28, 2022Updated 4 years ago
- Code for the paper Alpha Zero in Continuous Action Space (A0C) (https://arxiv.org/pdf/1805.09613.pdf)☆15Jan 19, 2021Updated 5 years ago
- Forked from: https://github.com/carcamdou/cr_grasper☆10Dec 7, 2020Updated 5 years ago
- Semantic-Aware Fine-Grained Correspondence, at ECCV 2022 (Oral)☆14Oct 29, 2022Updated 3 years ago
- 【云顶之弈小帮手】TFT-Helper☆15Jan 6, 2020Updated 6 years ago
- ACL'2025: SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs. and preprint: SoftCoT++: Test-Time Scaling with Soft Chain-of…☆96May 30, 2025Updated last year
- [NeurIPS 2024] Can Language Models Learn to Skip Steps?☆22Jan 25, 2025Updated last year
- ComiRec reproduced using pytorch 1.7.1 and python 3.7☆21Sep 7, 2022Updated 4 years ago
- NEO is a LLM inference engine built to save the GPU memory crisis by CPU offloading☆99Jun 16, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆27Jul 22, 2022Updated 4 years ago
- ☆21Jun 9, 2025Updated last year
- [ICDE 2024] VDTuner - Automated Performance Tuning for Vector Data Management Systems (Vector Databases)☆34Apr 21, 2024Updated 2 years ago
- Benchmark AFLOW Data Sets for Machine Learning doi.org/10.1007/s40192-020-00174-4☆11Aug 29, 2020Updated 6 years ago
- 电动笛子!☆12Jul 13, 2024Updated 2 years ago
- ☆19May 9, 2025Updated last year
- [COLM 2025] Official code for "When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoni…☆15Oct 31, 2025Updated 11 months ago
- ☆13Jan 25, 2026Updated 8 months ago
- A compiler framework for eBPF programs☆23Oct 1, 2026Updated last week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A benchmarking framework for on-device AI☆21Aug 20, 2026Updated last month
- Code for paper "Hierarchically Decoupled Imitation for Morphological Transfer"☆17Mar 24, 2023Updated 3 years ago
- ☆12Apr 29, 2024Updated 2 years ago
- Source code of paper "Machine Learning for Load Balancing in the Linux Kernel"☆23Sep 21, 2020Updated 6 years ago
- [ASPLOS' 26] TetriServe: Efficiently Serving Mixed DiT Workloads☆20Mar 12, 2026Updated 6 months ago
- a starter-kit for jaynes, the cloud-agnostic launch library☆17Apr 1, 2026Updated 6 months ago
- The pytorch code for ICRA2020 paper 'CMTS: Conditional Multiple Trajectory Synthesizer for Generating Safety-critical Driving Scenarios'☆19May 18, 2020Updated 6 years ago
- ☆14Oct 3, 2024Updated 2 years ago
- [NeurIPS 2022 Spotlight] Improving 3D-aware Image Synthesis with A Geometry-aware Discriminator☆30Oct 3, 2022Updated 4 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Confidence Regulation Neurons in Language Models (NeurIPS 2024)☆16Feb 1, 2025Updated last year
- Dotfile management with bare git☆22Updated this week
- High-performance LLM inference based on our optimized version of FastTransfomer☆123Dec 14, 2023Updated 2 years ago
- [AAAI 2025] Neural-Symbolic Collaborative Distillation: Advancing Small Language Models for Complex Reasoning Tasks☆13Jun 19, 2025Updated last year
- This is the repository of the EnviroDetaNet☆14Sep 3, 2024Updated 2 years ago
- A PyTorch implementation of Generative Flows with Matrix Exponential.☆12Sep 22, 2024Updated 2 years ago
- "FiD-ICL: A Fusion-in-Decoder Approach for Efficient In-Context Learning" (ACL 2023)☆15Jul 24, 2023Updated 3 years ago