PolarEngine: vLLM plugin for PolarQuant quantized LLM inference — 75% FP16 speed at 2.3x less VRAM
☆35Apr 13, 2026Updated 4 months ago
Alternatives and similar repositories for polarengine-vllm
Users that are interested in polarengine-vllm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- EOQ: Entropy-Optimal Quantization for LLMs. 11-41% smaller than GGUF Q4_K_M with near-FP16 perplexity.☆46Aug 15, 2026Updated last week
- ☆15Apr 26, 2025Updated last year
- Talk to your Claude Code by voice — a real voice call with your terminal agent. Local whisper STT, edge-tts voice, no API key.☆36Jul 4, 2026Updated last month
- Structured Chain-of-Thought☆220May 16, 2026Updated 3 months ago
- AI Agents Workshop with Red Hat AI☆13Feb 26, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆20Apr 27, 2025Updated last year
- Video2Video Framework for ComfyUI☆63Aug 12, 2024Updated 2 years ago
- ☆19Sep 4, 2024Updated last year
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆231Updated this week
- ComfyUI nodes for editing background of images/videos with CUDA acceleration support.☆22Dec 31, 2024Updated last year
- a browser gui for nvidia smi☆21Mar 17, 2025Updated last year
- ☆13May 23, 2024Updated 2 years ago
- Unofficial implementation of DreamTalk in ComfyUI☆12Aug 15, 2024Updated 2 years ago
- Nodes to level up your workflows performance and streamline specific functions.☆11Aug 19, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆12Nov 26, 2024Updated last year
- marlin_v100 是一个从 vLLM 主树中提取出来的最小 Marlin 独立开发工作区,聚焦于 Marlin dense 与 Marlin MoE 的源码开发、最小构建和轻量验证。它保留了核心 CUDA/C++ 实现、最小 Python 薄封装、生成器测试与主树回写…☆23Jul 2, 2026Updated last month
- Demo showing you calling ComfyUI API with ComfyDeploy☆20Aug 29, 2025Updated 11 months ago
- Ibis is a Hands-Free Interactive Web Page. Using the latest generative AI, it can be Any Page.☆21Oct 30, 2024Updated last year
- Local benchmarking UI for LLMs and AI agents☆23Apr 13, 2026Updated 4 months ago
- Running a 32 GB AI model on 28 GB of memory — MoE expert streaming from NVMe SSD on Windows☆26Mar 31, 2026Updated 4 months ago
- Voice synthesis library for Text-to-Speech applications (HTS Engine rewrite in Rust language)☆13Updated this week
- ☆13Dec 20, 2025Updated 8 months ago
- ☆395Apr 16, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Hermes Agent plugin: transparent per-token cost accounting for local LLM inference rigs☆24Aug 1, 2026Updated 3 weeks ago
- PyTTI (pytti-tools) is a CLIP-guided image and animation synthesis toolkit.☆15Apr 24, 2026Updated 4 months ago
- FastAPI wrapper around original Vibevoice 1.5B and 7B models, with support for AWQ4 quant☆33Jun 22, 2026Updated 2 months ago
- ☆10Oct 24, 2024Updated last year
- ☆93Jul 7, 2025Updated last year
- 🎙️ VibeVoice FastAPI - Multi-Speaker TTS API☆33Aug 27, 2025Updated 11 months ago
- LMDB Adapter for gunDB☆14Dec 8, 2022Updated 3 years ago
- Own your AI, search the web with it🌐😎☆97Jan 14, 2025Updated last year
- MCP as a Judge is a behavioral MCP that strengthens AI coding assistants by requiring explicit LLM evaluations☆17Dec 15, 2025Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- WebRTC plugin for IE☆12Aug 22, 2017Updated 9 years ago
- ☆25Jul 30, 2026Updated 3 weeks ago
- Experimental llama.cpp fork for inference research and development☆779Updated this week
- PyTorch -> ONNX☆17Oct 18, 2025Updated 10 months ago
- Local runner for Microsoft VibeVoice Realtime TTS Fully compatible with Open-Webui Plug and Play. OpenAI api endpoint .Run the Colab note…☆44Jul 20, 2026Updated last month
- Python package for extractive NLP using the OpenAI API☆17Aug 28, 2024Updated last year
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,043Apr 23, 2026Updated 4 months ago