A high-throughput and memory-efficient inference and serving engine for LLMs
☆39Jun 24, 2026Updated last month
Alternatives and similar repositories for vllm
Users that are interested in vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆74Jul 23, 2026Updated 3 weeks ago
- Official code for Guiding Language Model Math Reasoning with Planning Tokens☆19Feb 29, 2024Updated 2 years ago
- DeepSeek-V4 Lecture☆27Updated this week
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆23Jan 25, 2026Updated 6 months ago
- A proxy for Google Bard LLM☆10Nov 2, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- LLM graph-RAG SQL generator for large databases with poor documentation☆19Sep 12, 2024Updated last year
- A retrieval augmented sequence modeling toolkit implemented based on Fairseq☆29Mar 3, 2023Updated 3 years ago
- Code for paper: [ICLR2025 Oral] FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference☆170Oct 13, 2025Updated 10 months ago
- React app for inspecting, building and debugging with the Realtime API☆11Nov 5, 2024Updated last year
- MacPilot CLI is a mcp server, It provides a collection of system tools that allow AI assistants to perform various operations on macOS sy…☆16May 10, 2025Updated last year
- ☆32May 26, 2024Updated 2 years ago
- tabular q learning for trading☆12Dec 10, 2018Updated 7 years ago
- (AAAI24 oral) Implementation of RPPO(Risk-sensitive PPO) and RPBT(Population-based self-play with RPPO)☆12May 22, 2023Updated 3 years ago
- Live stock sentiment dashboard of Dow Jones stock, showing the sentiment, stocks, industries and their respective allocation in the Dow J…☆12Dec 19, 2025Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆75Mar 26, 2025Updated last year
- Julia implementation of flash-attention operation for neural networks.☆11May 31, 2023Updated 3 years ago
- Community maintained hardware plugin for vLLM on Ascend☆2,611Updated this week
- A simple C Thread pool implementation☆13Apr 10, 2020Updated 6 years ago
- NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning Tasks☆20May 10, 2022Updated 4 years ago
- 数据标注工具label-studio汉化版☆17Jul 16, 2021Updated 5 years ago
- ☆79Dec 15, 2023Updated 2 years ago
- Sparse symmetric indefinite solver implemented with a runtime system☆13May 11, 2020Updated 6 years ago
- ☆19Feb 25, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A graph database based on Python,一个轻量级图数据库☆19Dec 13, 2020Updated 5 years ago
- An MPI wrapper for the pytorch tensor library that is automatically differentiable☆10Mar 27, 2023Updated 3 years ago
- Bagua tutorials.☆13Sep 4, 2022Updated 3 years ago
- FTRL-Proximal Online Learning Algorithm☆15May 22, 2017Updated 9 years ago
- ☆34Jul 27, 2026Updated 2 weeks ago
- allowing R users to work with dlib through Rcpp☆13Apr 11, 2018Updated 8 years ago
- Distributed SDDMM Kernel☆13Jul 8, 2022Updated 4 years ago
- Parse Networking PDU , *.pcap *.pcapng file And Networing knowledge.☆14May 20, 2022Updated 4 years ago
- ROUGE L metric implementation using tensorflow ops☆12Sep 17, 2018Updated 7 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- We propose the Text-to-CQL task and provide the dataset.☆35Jun 26, 2023Updated 3 years ago
- Google AI 2018 BERT pytorch implementation☆13Oct 22, 2018Updated 7 years ago
- aigc evals☆10Dec 2, 2023Updated 2 years ago
- A high-performance, generic MPMC queue in Go that uses buffer swapping to minimize contention and optimize throughput.☆21Aug 19, 2025Updated 11 months ago
- LaTeX Examples Document Source☆11Apr 9, 2024Updated 2 years ago
- vLLM adapter for a TGIS-compatible gRPC server.☆57Updated this week
- 使用SwiftUI开发的ChatGPT聊天APP☆16May 1, 2023Updated 3 years ago