☆18Feb 24, 2026Updated 5 months ago
Alternatives and similar repositories for vllm-learn
Users that are interested in vllm-learn are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A fault-tolerant RDMA-based disaggregated key-value store with 1-RTT UPDATEs and GETs thanks to the SWARM replication protocol☆14Sep 25, 2024Updated last year
- ☆16Feb 8, 2024Updated 2 years ago
- ☆16Apr 13, 2024Updated 2 years ago
- Mirror of the Xen Repository (PRs not accepted see: http://wiki.xenproject.org/wiki/Submitting_Xen_Project_Patches)☆17Sep 12, 2017Updated 8 years ago
- Source code for the paper "LongGenBench: Long-context Generation Benchmark"☆24Oct 8, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Official code and resources for the paper "EXIT: Context-Aware Extractive Compression for Enhancing Retrieval-Augmented Generation."☆26Jul 15, 2026Updated last month
- ☆12Mar 13, 2023Updated 3 years ago
- some hexagon intrinsic examples based on Qualcomm Hexagon☆17Mar 7, 2025Updated last year
- Windows paravirtualized☆27Sep 5, 2025Updated 11 months ago
- ☆11Jun 6, 2023Updated 3 years ago
- A C++ port of karpathy/micrograd, a tiny scalar-valued autograd engine and a neural net library☆13Nov 24, 2023Updated 2 years ago
- Artifact evaluation repo for EuroSys'24.☆29Nov 7, 2023Updated 2 years ago
- A c++ hash map/table which utilizes simd (specifically Intel x86 SSE/AVX)☆12Apr 30, 2019Updated 7 years ago
- ☆26May 30, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- KV cache compression for high-throughput LLM inference☆159Feb 5, 2025Updated last year
- ☆17Jan 3, 2025Updated last year
- OneFlow Serving☆20Apr 10, 2025Updated last year
- ☆33Sep 29, 2021Updated 4 years ago
- Multiple GEMM operators are constructed with cutlass to support LLM inference.☆20Aug 3, 2025Updated last year
- LazyLog: A New Shared Log Abstraction for Low-Latency Applications☆47Apr 28, 2025Updated last year
- A torch compile backend for multi-targets☆51May 27, 2026Updated 2 months ago
- An IR for efficiently simulating distributed ML computation.☆33Jan 13, 2024Updated 2 years ago
- Pytorch implementation of our paper accepted by ICML 2024 -- CaM: Cache Merging for Memory-efficient LLMs Inference☆50Jun 19, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- TLLM_QMM strips the implementation of quantized kernels of Nvidia's TensorRT-LLM, removing NVInfer dependency and exposes ease of use Pyt…☆16Jul 5, 2024Updated 2 years ago
- ☆47Nov 25, 2024Updated last year
- ☆27Jun 8, 2026Updated 2 months ago
- Homework of CMU 10-414/714: Deep Learning Systems (https://dlsyscourse.org/)☆15Mar 21, 2024Updated 2 years ago
- This repository contains some sentiment analysis models and sequence tagging models, including BiLSTM, TextCNN, BERT for both tasks. All …☆13Feb 1, 2023Updated 3 years ago
- Parallel selection on GPUs☆15Mar 23, 2021Updated 5 years ago
- Converting Chinese sentences into pinyin sequences, implemented in C++, very fast and easy to deploy.☆23Jan 5, 2026Updated 7 months ago
- [ICLR2025] Code and data for paper: Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasonin…☆45Mar 10, 2025Updated last year
- ☆19Apr 6, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆27Feb 17, 2025Updated last year
- Analysis for the traces from byteprofile☆32Nov 21, 2023Updated 2 years ago
- A small diffusion model for text-to-image generation☆24Mar 10, 2025Updated last year
- Easy control for Key-Value Constrained Generative LLM Inference(https://arxiv.org/abs/2402.06262)☆62Feb 13, 2024Updated 2 years ago
- Möbius Transformation for Fast Inner Product Search on Graph☆23Jun 3, 2021Updated 5 years ago
- Keyformer proposes KV Cache reduction through key tokens identification and without the need for fine-tuning☆58Mar 26, 2024Updated 2 years ago
- 这是为希望学习FAISS向量数据库的同学准备的全面入门指导,帮助你快速建立相关概念,更好地阅读官方文档。☆37Nov 6, 2025Updated 9 months ago