☆560Oct 9, 2026Updated this week
Alternatives and similar repositories for llm-scaler
Users that are interested in llm-scaler are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Inference engine for Intel devices. Serve LLMs, VLMs, Whisper, Kokoro-TTS, Embedding and Rerank models over OpenAI endpoints.☆538Updated this week
- ☆197Updated this week
- The vLLM XPU kernels for Intel GPU☆75Updated this week
- ☆61Sep 14, 2026Updated 3 weeks ago
- SGLang kernel library for Intel XPU☆36Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- AI PC starter app for doing AI image creation, image stylizing, and chatbot on a PC powered by an Intel® Arc™ GPU.☆1,000Oct 1, 2026Updated last week
- ☆115Updated this week
- A simple and effective post training quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的后训练量化工具包☆1,630Updated this week
- Run Generative AI models with simple C++/Python API and using OpenVINO Runtime☆600Updated this week
- 🤗 Optimum Intel: Accelerate inference with Intel optimization tools☆622Updated this week
- llama-benchy - llama-bench style benchmarking tool for all backends☆751Jul 10, 2026Updated 2 months ago
- The AI PC Application Installer provides a unified way to set up Intel AI PC development environments.☆45Updated this week
- Cache-DiT Node for Comfyui☆308Aug 4, 2026Updated 2 months ago
- This repository contains Dockerfiles, scripts, yaml files, Helm charts, etc. used to scale out AI containers with versions of TensorFlow …☆79Sep 8, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆158Updated this week
- With OpenVINO Test Drive, users can run large language models (LLMs) and models trained by Intel Geti on their devices, including AI PCs …☆41Sep 1, 2026Updated last month
- Developer kits reference setup scripts for various kinds of Intel platforms and GPUs☆54Updated this week
- Community maintained hardware plugin for vLLM on Intel Gaudi☆61Updated this week
- OpenVINO™ is an open source toolkit for optimizing and deploying AI inference☆10,974Updated this week
- Intel® NPU (Neural Processing Unit) Driver☆465Updated this week
- ☆29Jun 10, 2026Updated 4 months ago
- ☆20Sep 28, 2026Updated last week
- Open-source cross-modal and multimodal prompt injection test suite. 250,000+ attack payloads across text, image, document, and audio moda…☆76Jul 22, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Edge Insights for Vision (eiv) is a package that helps to auto install Intel® GPU drivers and setup environment for Inference application…☆22Sep 29, 2025Updated last year
- Intel® AI Builder (SuperClaw, SuperBuilder)☆257Updated this week
- A Python package for extending the official PyTorch that can easily obtain performance on Intel platform☆2,011Mar 30, 2026Updated 6 months ago
- ☆18Oct 2, 2026Updated last week
- Messy repo filled with messy tests about hardware and LLMs. Built for me, public for you.☆47Sep 27, 2026Updated last week
- Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https…☆5,851Updated this week
- OpenVINO LLM Benchmark☆11Dec 7, 2023Updated 2 years ago
- Large Language Model Text Generation Inference on Habana Gaudi☆34Mar 20, 2025Updated last year
- Explore our open source AI portfolio! Develop, train, and deploy your AI solutions with performance- and productivity-optimized tools fro…☆79Mar 27, 2026Updated 6 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Libraries, microservices, tools, and other reference software, supporting development of performance-optimized Edge AI applications.☆170Updated this week
- ONNX Runtime: cross-platform, high performance scoring engine for ML models☆92Updated this week
- AI plays Doom — pit Vision Language Models against demons and each other. Solo scenarios, deathmatch arena, 1-4 agents with any OpenAI-co…☆21Mar 12, 2026Updated 6 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,288Updated this week
- FORK of VLLM for AMD MI25/50/60. A high-throughput and memory-efficient inference and serving engine for LLMs☆70May 4, 2025Updated last year
- Samples running deep learning models on Intel GPU Arc A770☆15Jul 4, 2024Updated 2 years ago
- Enable true multi gpu capability in Comfy UI using XDiT XFuser and FSDP managed by Ray☆465Updated this week