A high-throughput and memory-efficient inference and serving engine for LLMs
☆90Jul 27, 2026Updated 2 weeks ago
Alternatives and similar repositories for vllm-fork
Users that are interested in vllm-fork are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16Jun 4, 2026Updated 2 months ago
- ☆19Jul 13, 2026Updated last month
- Easy and lightning fast training of 🤗 Transformers on Habana Gaudi processor (HPU)☆212Updated this week
- Community maintained hardware plugin for vLLM on Intel Gaudi☆51Updated this week
- Reference models for Intel(R) Gaudi(R) AI Accelerator☆172Jan 8, 2026Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- SynapseAI Core is a reference implementation of the SynapseAI API running on Habana Gaudi☆46Feb 3, 2025Updated last year
- Provides the examples to write and build Habana custom kernels using the HabanaTools☆26Apr 15, 2025Updated last year
- PM Workshop China☆10Apr 11, 2019Updated 7 years ago
- A comprehensive, enterprise-ready toolkit for deploying Agentic AI systems on Intel® Xeon processors and Intel accelerators. Built for or…☆21Updated this week
- The vLLM XPU kernels for Intel GPU☆60Updated this week
- This is a fork of SGLang for hip-attention integration. Please refer to hip-attention for detail.☆18Mar 31, 2026Updated 4 months ago
- SYCL* Templates for Linear Algebra (SYCL*TLA) - SYCL based CUTLASS implementation for Intel GPUs☆80Updated this week
- GenAI components at micro-service level; GenAI service composer to create mega-service☆199Updated this week
- ☆26Oct 9, 2025Updated 10 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Tools and pipelines for automated LLM performance evaluation☆15May 20, 2026Updated 2 months ago
- Setup and Installation Instructions for Habana binaries, docker image creation☆28Jul 10, 2026Updated last month
- ☆17Dec 29, 2025Updated 7 months ago
- ☆18May 29, 2026Updated 2 months ago
- Intel Gaudi's Megatron DeepSpeed Large Language Models for training☆18Dec 19, 2024Updated last year
- QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference☆123Mar 6, 2024Updated 2 years ago
- Intel® AI for Enterprise Inference optimizes AI inference services on Intel hardware using Kubernetes Orchestration. It automates LLM mod…☆45Updated this week
- ☆102Updated this week
- AI quickstart that provides interactive dashboard to analyze AI Model Performance as well as Openshift metrics collected from Prometheus☆26Jun 9, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Intel® Extension for DeepSpeed* is an extension to DeepSpeed that brings feature support with SYCL kernels on Intel GPU(XPU) device. Note…☆65May 27, 2026Updated 2 months ago
- SGLang kernel library for Intel XPU☆29Updated this week
- vLLM: A high-throughput and memory-efficient inference and serving engine for LLMs☆95Updated this week
- OpenVINO LLM Benchmark☆11Dec 7, 2023Updated 2 years ago
- Official Code Repository for the paper "KALA: Knowledge-Augmented Language Model Adaptation" (NAACL 2022)☆35Oct 17, 2023Updated 2 years ago
- The Prometheus monitoring system and time series database.☆15Aug 6, 2026Updated last week
- ☆19Jul 24, 2025Updated last year
- Cloud Native Benchmarking of Foundation Models☆46Jul 31, 2025Updated last year
- Ongoing research training transformer models at scale☆43Updated this week
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Mini-Engine Demonstration of Combining XeSS with VRS Tier 2.☆14Jan 26, 2026Updated 6 months ago
- 🚀 Collection of components for development, training, tuning, and inference of foundation models leveraging PyTorch native components.☆233Aug 4, 2026Updated last week
- Benchmarking for multiple AWS S3 libraries.☆16Updated this week
- Benchmark Suite Invocation Scripting☆11Mar 16, 2022Updated 4 years ago
- The kernel module management operator builds, signs and loads kernel modules on OpenShift.☆31Updated this week
- Official Code Repository for the paper "Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-intensive Tasks…☆44Nov 24, 2024Updated last year
- SPDK fork of nvme-cli. No longer supported - use standard nvme-cli with SPDK nvme CUSE instead. See https://spdk.io/doc/nvme.html#nvme_…☆15Apr 10, 2024Updated 2 years ago