Why is LLM inference slow — and how do you make it fast? A hands-on, first-principles course: roofline → KV cache → quantization → parallelism → vLLM/SGLang, with GPU labs on open models.
☆20Aug 11, 2026Updated 2 weeks ago
Alternatives and similar repositories for Efficient-LLM-Inference-Serving-Systems
Users that are interested in Efficient-LLM-Inference-Serving-Systems are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆10Dec 15, 2024Updated last year
- Nvim plugin for reviewing PRs with side by side diffs, inline comments, and persistent review state.☆16Apr 3, 2026Updated 4 months ago
- ☆10May 29, 2024Updated 2 years ago
- ☆13Aug 17, 2020Updated 6 years ago
- CastleHill: Separable Causal Diffusion / Varitaion Flow Maps for LTX-2 long-form video generation☆15May 19, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Deriving steepest descent convergence bounds and hyperparameter scaling laws in machine learning optimization from first principles, form…☆17Apr 11, 2026Updated 4 months ago
- ☆11Feb 22, 2025Updated last year
- Full End-to-End examples showing how to use First-gen Gaudi and Gaudi2 in common use cases☆13Dec 2, 2024Updated last year
- Repository for NLP project. Name to be changed when we decide on a project☆17Apr 19, 2022Updated 4 years ago
- ☆15Apr 14, 2025Updated last year
- VidKV: Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models☆25Mar 26, 2025Updated last year
- Iterate fast on your RAG pipelines☆24Jun 21, 2025Updated last year
- ☆18Nov 22, 2022Updated 3 years ago
- ☆30Sep 3, 2021Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Repository for ACM India Summer School on Generative AI for Text☆13Jul 11, 2024Updated 2 years ago
- Replication package for evaluation of code generation metrics☆17Nov 24, 2025Updated 9 months ago
- Code for "Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models".☆24Mar 16, 2026Updated 5 months ago
- ☆37Jul 5, 2024Updated 2 years ago
- ☆15Aug 26, 2023Updated 3 years ago
- Example of using the SAINT architecture in fastai☆12Aug 25, 2021Updated 5 years ago
- Use the Google Cloud Speech API to transcribe audio files from a podcast.☆20May 17, 2017Updated 9 years ago
- Triton implementation of GPT/LLAMA☆22Aug 28, 2024Updated 2 years ago
- Image-to-text translation of chemical molecule structures with deep learning (top-5% Kaggle solution)☆15Sep 4, 2022Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Forecastable Component Analysis (ForeCA) in Python☆38Nov 21, 2025Updated 9 months ago
- This repository provides the official implementation of QSVD, a method for efficient low-rank approximation that unifies Query-Key-Value …☆28May 16, 2026Updated 3 months ago
- Utilizing nbdev in Google Colaboratory☆14Apr 12, 2023Updated 3 years ago
- A curated list of resources dedicated to Code-mixed Natural Language Processing (NLP).☆18Jun 23, 2026Updated 2 months ago
- DREN:Deep Rotation Equivirant Network☆16Mar 24, 2019Updated 7 years ago
- ☆11Apr 27, 2019Updated 7 years ago
- republish livox raw message to standard pointcloud2☆38Jul 9, 2023Updated 3 years ago
- Brain-Box is a web application that allows students to organize and manage their study materials, including subjects, chapters, notes, an…☆27Feb 20, 2024Updated 2 years ago
- Get images from MangaPanda, MangaReader or MangaFox and save it for offline reading.☆11Jul 5, 2016Updated 10 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code for "Simulated Multiple Reference Training Improves Low-Resource Machine Translation"☆15Dec 1, 2020Updated 5 years ago
- ☆28Aug 22, 2026Updated last week
- treemind interprets tree models☆40Apr 9, 2026Updated 4 months ago
- A LLM Agent with Langchain/Langgraph helps to analyze CV, look relevant jobs via API, and write a cover letter according to it☆61May 1, 2024Updated 2 years ago
- Text Normalization utilities for normalizing text for TTS☆26Mar 4, 2026Updated 5 months ago
- Implementation of <Orthogonal Model Merging>☆35Aug 5, 2026Updated 3 weeks ago
- Public-facing codebase accompanying: "Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL"☆37Feb 6, 2026Updated 6 months ago