A lightweight and fast LLM serving framework
☆15Mar 5, 2026Updated 4 months ago
Alternatives and similar repositories for LMServe
Users that are interested in LMServe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [IEEE CAL 2025] Accelerating Page Migrations in Operating Systems with Intel DSA☆16Nov 20, 2024Updated last year
- [USENIX ATC 2021] Exploring the Design Space of Page Management for Multi-Tiered Memory Systems☆49Mar 31, 2022Updated 4 years ago
- [ACM EuroSys 2023] Fast and Efficient Model Serving Using Multi-GPUs with Direct-Host-Access☆56Aug 6, 2025Updated 11 months ago
- ☆27Aug 19, 2022Updated 3 years ago
- This is the respository that holds the artifacts of ASPLOS'25 -- M5: Mastering Page Migration and Memory Management for CXL-based Tiered …☆17Apr 1, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆43Sep 3, 2025Updated 10 months ago
- Artifact package for CBMM paper (ATC'22)☆11Jun 5, 2022Updated 4 years ago
- Heterogeneous Memory Software Development Kit☆101Mar 9, 2026Updated 4 months ago
- ☆15Jul 10, 2026Updated last week
- Official code repository for "CoVA: Exploiting Compressed-Domain Analysis to Accelerate Video Analytics [USENIX ATC 22]"☆18Sep 19, 2024Updated last year
- OSDI'24 Nomad implementation☆55Aug 1, 2025Updated 11 months ago
- On-demand-fork☆33Mar 28, 2023Updated 3 years ago
- ☆21Jul 13, 2026Updated last week
- ☆30May 20, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆24Dec 6, 2025Updated 7 months ago
- ISCA-2025☆25Mar 3, 2026Updated 4 months ago
- A comprehensive repository for Compute Express Link (CXL) resources: covering research papers, specifications, simulation/emulation tools…☆26Feb 24, 2026Updated 4 months ago
- SNS Hashtag Offilne Event Managing Platform☆12Feb 11, 2022Updated 4 years ago
- ☆10Jun 1, 2024Updated 2 years ago
- Repo to replay Qwen trace☆31Jan 9, 2026Updated 6 months ago
- Tiered memory management☆90Sep 1, 2025Updated 10 months ago
- Tiered Memory Management Beyond Hotness (OSDI'25)☆37Jul 31, 2025Updated 11 months ago
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]☆12Nov 8, 2024Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Latr: Lazy Translation Coherence - ASPLOS'18☆15Nov 15, 2021Updated 4 years ago
- ☆10Sep 15, 2023Updated 2 years ago
- Various Sorting Algorithms with golang☆11Mar 22, 2026Updated 3 months ago
- The source code of INFless,a native serverless platform for AI inference.☆46Oct 10, 2022Updated 3 years ago
- An easy and ready-to-go bootstrap for k8s installation and automatic cluster deployment!☆11Jan 5, 2025Updated last year
- A disaggregated memory orchestration system that virtualizes cluster wide memory to scale data intensive, large memory workloads in virtu…☆13Apr 26, 2019Updated 7 years ago
- ☆27Aug 31, 2023Updated 2 years ago
- Demo repository for all the different ways to do eBPF Tracing☆18Feb 9, 2026Updated 5 months ago
- SwarmIO is an SSD emulation framework for next-generation GPU-centric storage systems research☆52May 24, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- MetaAttention: A Unified and Performant Attention Framework Across Hardware Backends(PPoPP'26)☆16Dec 31, 2025Updated 6 months ago
- Kernel repo of "Nimble Page Management for Tiered Memory Systems" in ASPLOS 2019☆46Aug 2, 2022Updated 3 years ago
- ☆16Dec 4, 2025Updated 7 months ago
- Demo Repository for eBPF XDP Unit Test☆12Oct 24, 2024Updated last year
- Proteus: A High-Throughput Inference-Serving System with Accuracy Scaling☆13Mar 7, 2024Updated 2 years ago
- Prefix-Aware Attention for LLM Decoding☆41May 26, 2026Updated last month
- Artifact for Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization☆17May 9, 2025Updated last year