☆511Aug 29, 2026Updated this week
Alternatives and similar repositories for llm-scaler
Users that are interested in llm-scaler are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Inference engine for Intel devices. Serve LLMs, VLMs, Whisper, Kokoro-TTS, Embedding and Rerank models over OpenAI endpoints.☆513Aug 22, 2026Updated last week
- ☆191Aug 20, 2026Updated last week
- The vLLM XPU kernels for Intel GPU☆65Updated this week
- ☆58Jul 16, 2026Updated last month
- ☆106Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- AI PC starter app for doing AI image creation, image stylizing, and chatbot on a PC powered by an Intel® Arc™ GPU.☆968Updated this week
- A SOTA quantization toolkit for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support a…☆1,594Updated this week
- Run Generative AI models with simple C++/Python API and using OpenVINO Runtime☆582Updated this week
- 🤗 Optimum Intel: Accelerate inference with Intel optimization tools☆615Updated this week
- llama-benchy - llama-bench style benchmarking tool for all backends☆682Jul 10, 2026Updated last month
- This repository contains Dockerfiles, scripts, yaml files, Helm charts, etc. used to scale out AI containers with versions of TensorFlow …☆79May 27, 2026Updated 3 months ago
- ☆154Updated this week
- With OpenVINO Test Drive, users can run large language models (LLMs) and models trained by Intel Geti on their devices, including AI PCs …☆41Mar 12, 2026Updated 5 months ago
- Developer kits reference setup scripts for various kinds of Intel platforms and GPUs☆50Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆30Updated this week
- Community maintained hardware plugin for vLLM on Intel Gaudi☆52Updated this week
- Intel® NPU (Neural Processing Unit) Driver☆450Aug 7, 2026Updated 3 weeks ago
- ☆28Jun 10, 2026Updated 2 months ago
- ☆20Updated this week
- Intel® AI Builder (SuperClaw, SuperBuilder)☆236Updated this week
- Open-source cross-modal and multimodal prompt injection test suite. 250,000+ attack payloads across text, image, document, and audio moda…☆70Jul 22, 2026Updated last month
- A Python package for extending the official PyTorch that can easily obtain performance on Intel platform☆2,012Mar 30, 2026Updated 5 months ago
- ☆18Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Fully automated installation scripts for ComfyUI optimized for Intel Arc GPUs (A-Series) and Intel Core Ultra iGPUs with XPU backend, Tri…☆159Feb 10, 2026Updated 6 months ago
- Messy repo filled with messy tests about hardware and LLMs. Built for me, public for you.☆49Aug 17, 2026Updated last week
- Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https…☆5,538Updated this week
- OpenVINO LLM Benchmark☆11Dec 7, 2023Updated 2 years ago
- Libraries, microservices, tools, and other reference software, supporting development of performance-optimized Edge AI applications.☆156Updated this week
- Large Language Model Text Generation Inference on Habana Gaudi☆34Mar 20, 2025Updated last year
- Explore our open source AI portfolio! Develop, train, and deploy your AI solutions with performance- and productivity-optimized tools fro…☆78Mar 27, 2026Updated 5 months ago
- AI plays Doom — pit Vision Language Models against demons and each other. Solo scenarios, deathmatch arena, 1-4 agents with any OpenAI-co…☆20Mar 12, 2026Updated 5 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,148Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- FORK of VLLM for AMD MI25/50/60. A high-throughput and memory-efficient inference and serving engine for LLMs☆71May 4, 2025Updated last year
- Enable true multi gpu capability in Comfy UI using XDiT XFuser and FSDP managed by Ray☆418Aug 23, 2026Updated last week
- Generate a llama-quantize command to copy the quantization parameters of any GGUF☆36Apr 20, 2026Updated 4 months ago
- ☆1,214Updated this week
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,520Updated this week
- Make use of Intel Arc Series GPU to Run Ollama, StableDiffusion, Whisper and Open WebUI, for image generation, speech recognition and int…☆404Aug 7, 2026Updated 3 weeks ago
- Compressed KV cache as cross-backend wire format for Metal + CUDA split inference over Thunderbolt 5☆16Apr 14, 2026Updated 4 months ago