☆58Jul 16, 2026Updated last month
Alternatives and similar repositories for vllm-openvino
Users that are interested in vllm-openvino are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- OpenVINO Tokenizers extension☆56Updated this week
- Tools for easier OpenVINO development/debugging☆10Jul 16, 2025Updated last year
- Run Generative AI models with simple C++/Python API and using OpenVINO Runtime☆582Updated this week
- Community maintained hardware plugin for vLLM on Intel Gaudi☆52Updated this week
- SGLang kernel library for Intel XPU☆31Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- OpenVINO Intel NPU Compiler☆96Aug 24, 2026Updated last week
- Inference engine for Intel devices. Serve LLMs, VLMs, Whisper, Kokoro-TTS, Embedding and Rerank models over OpenAI endpoints.☆513Aug 22, 2026Updated last week
- This repository contains Dockerfiles, scripts, yaml files, Helm charts, etc. used to scale out AI containers with versions of TensorFlow …☆79May 27, 2026Updated 3 months ago
- 🤗 Optimum Intel: Accelerate inference with Intel optimization tools☆615Updated this week
- A lightweight triton-based General Matrix Multiplication (GEMM) library.☆67Jul 21, 2026Updated last month
- OpenAI Triton backend for Intel® GPUs☆267Updated this week
- Run a 35B MoE model at 10+ tok/s on a $600 Mac mini. Pure C/Metal inference engine streaming experts from SSD on Apple Silicon☆19Apr 20, 2026Updated 4 months ago
- Add genai backend for ollama to run generative AI models using OpenVINO Runtime.☆30Apr 16, 2026Updated 4 months ago
- ☆15Apr 11, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A Gradio Web UI for running local LLM on Intel GPU (e.g., local PC with iGPU, discrete GPU such as Arc, Flex and Max) using IPEX-LLM.☆17Aug 23, 2026Updated last week
- Mini-Engine Demonstration of Combining XeSS with VRS Tier 2.☆14Jan 26, 2026Updated 7 months ago
- PM Workshop China☆10Apr 11, 2019Updated 7 years ago
- ☆21Jul 3, 2024Updated 2 years ago
- Google DeepMind: Mixture of Depths Unofficial Implementation.☆12May 29, 2024Updated 2 years ago
- DLL注入工具☆13Nov 9, 2020Updated 5 years ago
- matmul using AMX instructions☆24May 7, 2024Updated 2 years ago
- SPDK fork of nvme-cli. No longer supported - use standard nvme-cli with SPDK nvme CUSE instead. See https://spdk.io/doc/nvme.html#nvme_…☆15Apr 10, 2024Updated 2 years ago
- GPU Functional Descriptor for memory access☆34May 24, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- LCM OpenVINO model converter☆24Mar 27, 2024Updated 2 years ago
- How to export PyTorch models with unsupported layers to ONNX and then to Intel OpenVINO☆28Feb 20, 2025Updated last year
- OpenGL 学习代码☆15Jun 25, 2023Updated 3 years ago
- This is a clone of an SVN repository at http://pagecache-mangagement.googlecode.com/svn/trunk. It had been cloned by http://svn2github.co…☆10May 23, 2013Updated 13 years ago
- An Awesome list of oneAPI projects☆164Aug 8, 2025Updated last year
- Pre-built components and code samples to help you build and deploy production-grade AI applications with the OpenVINO™ Toolkit from Intel☆217Jul 30, 2026Updated last month
- ☆15Jun 26, 2024Updated 2 years ago
- Unreal Engine 5 3D Platformer game prototype☆21May 27, 2024Updated 2 years ago
- OpenVINO™ is an open source toolkit for optimizing and deploying AI inference☆10,768Updated this week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Hexagon-MLIR is a compiler toolchain for compiling and executing AI kernels and models on Qualcomm Hexagon Neural Processing Units (NPUs)…☆205Updated this week
- Velocity And Luminance Adaptive Rasterization☆16Mar 31, 2023Updated 3 years ago
- code repo for GCR [FAST'26]☆16Mar 3, 2026Updated 5 months ago
- Intel® NPU (Neural Processing Unit) Driver☆450Aug 7, 2026Updated 3 weeks ago
- ☆18May 10, 2023Updated 3 years ago
- PilotFish harvests the free GPU cycles of cloud gaming with deep learning training☆14Jul 2, 2022Updated 4 years ago
- Service-aware KV-cache compression for bandwidth-efficient disaggregated LLM serving.☆22Jul 26, 2026Updated last month