A lightweight and fast LLM serving framework
☆15Mar 5, 2026Updated 5 months ago
Alternatives and similar repositories for LMServe
Users that are interested in LMServe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [IEEE CAL 2025] Accelerating Page Migrations in Operating Systems with Intel DSA☆16Nov 20, 2024Updated last year
- [USENIX ATC 2021] Exploring the Design Space of Page Management for Multi-Tiered Memory Systems☆49Mar 31, 2022Updated 4 years ago
- [ACM EuroSys 2023] Fast and Efficient Model Serving Using Multi-GPUs with Direct-Host-Access☆56Aug 6, 2025Updated last year
- ☆27Aug 19, 2022Updated 3 years ago
- This is the respository that holds the artifacts of ASPLOS'25 -- M5: Mastering Page Migration and Memory Management for CXL-based Tiered …☆17Apr 1, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆43Sep 3, 2025Updated 11 months ago
- Artifact package for CBMM paper (ATC'22)☆11Jun 5, 2022Updated 4 years ago
- Heterogeneous Memory Software Development Kit☆101Mar 9, 2026Updated 5 months ago
- ☆15Updated this week
- Official code repository for "CoVA: Exploiting Compressed-Domain Analysis to Accelerate Video Analytics [USENIX ATC 22]"☆18Sep 19, 2024Updated last year
- OSDI'24 Nomad implementation☆55Aug 1, 2025Updated last year
- On-demand-fork☆33Mar 28, 2023Updated 3 years ago
- ☆21Jul 13, 2026Updated 3 weeks ago
- ☆30May 20, 2026Updated 2 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ☆24Dec 6, 2025Updated 8 months ago
- ISCA-2025☆25Mar 3, 2026Updated 5 months ago
- ☆40Jun 2, 2026Updated 2 months ago
- A comprehensive repository for Compute Express Link (CXL) resources: covering research papers, specifications, simulation/emulation tools…☆26Feb 24, 2026Updated 5 months ago
- SNS Hashtag Offilne Event Managing Platform☆12Feb 11, 2022Updated 4 years ago
- Repo to replay Qwen trace☆31Jan 9, 2026Updated 7 months ago
- GDSC 고려대학교 nestJS 강의☆17Nov 29, 2023Updated 2 years ago
- DAMON user-space tool☆87Updated this week
- Tiered memory management☆91Sep 1, 2025Updated 11 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Tiered Memory Management Beyond Hotness (OSDI'25)☆37Jul 31, 2025Updated last year
- Latr: Lazy Translation Coherence - ASPLOS'18☆16Nov 15, 2021Updated 4 years ago
- Various Sorting Algorithms with golang☆11Mar 22, 2026Updated 4 months ago
- Medusa: Accelerating Serverless LLM Inference with Materialization [ASPLOS'25]☆12Nov 8, 2024Updated last year
- A Study of Database Performance Sensitivity to Experiment Settings☆11May 31, 2022Updated 4 years ago
- An easy and ready-to-go bootstrap for k8s installation and automatic cluster deployment!☆11Jan 5, 2025Updated last year
- The source code of INFless,a native serverless platform for AI inference.☆46Oct 10, 2022Updated 3 years ago
- A disaggregated memory orchestration system that virtualizes cluster wide memory to scale data intensive, large memory workloads in virtu…☆13Apr 26, 2019Updated 7 years ago
- ☆27Aug 31, 2023Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Demo repository for all the different ways to do eBPF Tracing☆19Feb 9, 2026Updated 6 months ago
- MetaAttention: A Unified and Performant Attention Framework Across Hardware Backends(PPoPP'26)☆17Updated this week
- SwarmIO is an SSD emulation framework for next-generation GPU-centric storage systems research☆54May 24, 2026Updated 2 months ago
- Proteus: A High-Throughput Inference-Serving System with Accuracy Scaling☆13Mar 7, 2024Updated 2 years ago
- Prefix-Aware Attention for LLM Decoding☆43May 26, 2026Updated 2 months ago
- ☆83Nov 16, 2020Updated 5 years ago
- Open-source of LazyDP published in ASPLOS-2024☆23May 5, 2024Updated 2 years ago