A lightweight and fast LLM serving framework
☆15Mar 5, 2026Updated 6 months ago
Alternatives and similar repositories for LMServe
Users that are interested in LMServe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [IEEE CAL 2025] Accelerating Page Migrations in Operating Systems with Intel DSA☆16Nov 20, 2024Updated last year
- [USENIX ATC 2021] Exploring the Design Space of Page Management for Multi-Tiered Memory Systems☆49Mar 31, 2022Updated 4 years ago
- [ACM EuroSys 2023] Fast and Efficient Model Serving Using Multi-GPUs with Direct-Host-Access☆56Aug 6, 2025Updated last year
- ☆27Aug 19, 2022Updated 4 years ago
- ☆44Sep 3, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Artifact package for CBMM paper (ATC'22)☆11Jun 5, 2022Updated 4 years ago
- Heterogeneous Memory Software Development Kit☆103Mar 9, 2026Updated 6 months ago
- ☆15Sep 10, 2026Updated last week
- Official code repository for "CoVA: Exploiting Compressed-Domain Analysis to Accelerate Video Analytics [USENIX ATC 22]"☆18Sep 19, 2024Updated last year
- OSDI'24 Nomad implementation☆55Aug 1, 2025Updated last year
- On-demand-fork☆33Mar 28, 2023Updated 3 years ago
- ☆22Jul 13, 2026Updated 2 months ago
- ☆31May 20, 2026Updated 3 months ago
- ☆24Dec 6, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ISCA-2025☆26Mar 3, 2026Updated 6 months ago
- ☆42Jun 2, 2026Updated 3 months ago
- A comprehensive repository for Compute Express Link (CXL) resources: covering research papers, specifications, simulation/emulation tools…☆27Feb 24, 2026Updated 6 months ago
- Repo to replay Qwen trace☆34Jan 9, 2026Updated 8 months ago
- GDSC 고려대학교 nestJS 강의☆17Nov 29, 2023Updated 2 years ago
- DAMON user-space tool☆90Updated this week
- Tiered memory management☆91Sep 1, 2025Updated last year
- Tiered Memory Management Beyond Hotness (OSDI'25)☆37Jul 31, 2025Updated last year
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]☆12Nov 8, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Various Sorting Algorithms with golang☆11Mar 22, 2026Updated 5 months ago
- A Study of Database Performance Sensitivity to Experiment Settings☆11May 31, 2022Updated 4 years ago
- An easy and ready-to-go bootstrap for k8s installation and automatic cluster deployment!☆11Aug 24, 2026Updated 3 weeks ago
- The source code of INFless,a native serverless platform for AI inference.☆46Oct 10, 2022Updated 3 years ago
- A disaggregated memory orchestration system that virtualizes cluster wide memory to scale data intensive, large memory workloads in virtu…☆13Apr 26, 2019Updated 7 years ago
- ☆27Aug 31, 2023Updated 3 years ago
- Demo repository for all the different ways to do eBPF Tracing☆19Feb 9, 2026Updated 7 months ago
- Kernel repo of "Nimble Page Management for Tiered Memory Systems" in ASPLOS 2019☆46Aug 2, 2022Updated 4 years ago
- MetaAttention: A Unified and Performant Attention Framework Across Hardware Backends(PPoPP'26)☆17Aug 6, 2026Updated last month
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆17Dec 4, 2025Updated 9 months ago
- SwarmIO is an SSD emulation framework for next-generation GPU-centric storage systems research☆58May 24, 2026Updated 3 months ago
- Demo Repository for eBPF XDP Unit Test☆12Oct 24, 2024Updated last year
- Proteus: A High-Throughput Inference-Serving System with Accuracy Scaling☆13Mar 7, 2024Updated 2 years ago
- Template repo for Python projects, especially those focusing on machine learning and/or deep learning.☆15Jan 14, 2026Updated 8 months ago
- ☆84Nov 16, 2020Updated 5 years ago
- Artifact for Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization☆18May 9, 2025Updated last year