☆458Aug 7, 2026Updated this week
Alternatives and similar repositories for llm-scaler
Users that are interested in llm-scaler are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Inference engine for Intel devices. Serve LLMs, VLMs, Whisper, Kokoro-TTS, Embedding and Rerank models over OpenAI endpoints.☆498Aug 3, 2026Updated last week
- ☆38May 29, 2026Updated 2 months ago
- ☆183Jul 25, 2026Updated 2 weeks ago
- The vLLM XPU kernels for Intel GPU☆60Updated this week
- ☆56Jul 16, 2026Updated 3 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- SGLang kernel library for Intel XPU☆29Updated this week
- ☆100Updated this week
- AI PC starter app for doing AI image creation, image stylizing, and chatbot on a PC powered by an Intel® Arc™ GPU.☆951Updated this week
- A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support…☆1,559Updated this week
- Run Generative AI models with simple C++/Python API and using OpenVINO Runtime☆567Updated this week
- 🤗 Optimum Intel: Accelerate inference with Intel optimization tools☆611Updated this week
- llama-benchy - llama-bench style benchmarking tool for all backends☆623Jul 10, 2026Updated last month
- Cache-DiT Node for Comfyui☆302Updated this week
- The AI PC Application Installer provides a unified way to set up Intel AI PC development environments.☆40Jul 23, 2026Updated 2 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- This repository contains Dockerfiles, scripts, yaml files, Helm charts, etc. used to scale out AI containers with versions of TensorFlow …☆79May 27, 2026Updated 2 months ago
- ☆153Updated this week
- With OpenVINO Test Drive, users can run large language models (LLMs) and models trained by Intel Geti on their devices, including AI PCs …☆39Mar 12, 2026Updated 4 months ago
- Developer kits reference setup scripts for various kinds of Intel platforms and GPUs☆50Updated this week
- Accelerate local LLM inference and finetuning (LLaMA, Mistral, ChatGLM, Qwen, DeepSeek, Mixtral, Gemma, Phi, MiniCPM, Qwen-VL, MiniCPM-V,…☆8,862Jan 28, 2026Updated 6 months ago
- vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60☆26Jul 2, 2026Updated last month
- Intel® NPU (Neural Processing Unit) Driver☆443Updated this week
- ☆25Jun 10, 2026Updated 2 months ago
- ☆19Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Intel® AI Builder (SuperClaw, SuperBuilder)☆212Updated this week
- A Python package for extending the official PyTorch that can easily obtain performance on Intel platform☆2,013Mar 30, 2026Updated 4 months ago
- ☆18Jul 23, 2026Updated 2 weeks ago
- Fully automated installation scripts for ComfyUI optimized for Intel Arc GPUs (A-Series) and Intel Core Ultra iGPUs with XPU backend, Tri…☆156Feb 10, 2026Updated 5 months ago
- Messy repo filled with messy tests about hardware and LLMs. Built for me, public for you.☆46Aug 2, 2026Updated last week
- Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https…☆5,298Updated this week
- OpenVINO LLM Benchmark☆11Dec 7, 2023Updated 2 years ago
- Performance optimized libraries, microservices, and tools to support the development of Edge AI applications.☆156Updated this week
- Large Language Model Text Generation Inference on Habana Gaudi☆34Mar 20, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Explore our open source AI portfolio! Develop, train, and deploy your AI solutions with performance- and productivity-optimized tools fro…☆78Mar 27, 2026Updated 4 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,021Updated this week
- Extracts (tangles) code fragments from Markdown documents.☆17Apr 10, 2026Updated 4 months ago
- Samples running deep learning models on Intel GPU Arc A770☆14Jul 4, 2024Updated 2 years ago
- Get aid from local LLMs right in your PowerShell☆16May 2, 2025Updated last year
- Enable true multi gpu capability in Comfy UI using XDiT XFuser and FSDP managed by Ray☆390Updated this week
- Generate a llama-quantize command to copy the quantization parameters of any GGUF☆36Apr 20, 2026Updated 3 months ago