The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with 100+ tok/s single-request decode with support of FP8 weight
☆459Jul 13, 2026Updated 2 weeks ago
Alternatives and similar repositories for vLLM-2080Ti-Definitive
Users that are interested in vLLM-2080Ti-Definitive are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆145Jun 13, 2026Updated last month
- vLLM fork for Tesla V100 (SM70) with AWQ 4-bit support, CUDA 12.8 build flow, and validated Qwen3.5 27B/35B deployment on multi-GPU V…☆557Jul 21, 2026Updated last week
- ☆16Apr 6, 2023Updated 3 years ago
- This project is specifically developed for V100, based on lmdeploy 0.12.1, and supports mainstream open-source models from Q4 2025 to Q1 …☆21Mar 18, 2026Updated 4 months ago
- Sage attention for turning.☆71Dec 29, 2025Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parall…☆100Updated this week
- ☆23Oct 12, 2022Updated 3 years ago
- 大语言模型工具集☆28Aug 1, 2025Updated 11 months ago
- A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations☆19,093Updated this week
- An extension utility for llama.cpp, used with 3090*2 + Strix Halo. llama.cpp的拓展小工具,自用于3090*2 + Strix Halo。☆282Updated this week
- fastllm是后端无依赖的高性能大模型推理库。同时支持张量并行推理稠密模型和混合模式推理MOE模型,任意10G以上显卡即可推理满血DeepSeek。双路9004/9005服务器+单显卡部署DeepSeek满血满精度原版模型,单并发20tps;INT4量化模型单并发30tp…☆4,873Updated this week
- KTransformers 一键部署脚本☆60Apr 18, 2025Updated last year
- One-KVM,结合廉价硬件和PiKVM,实现低成本远控方案。☆10Sep 6, 2024Updated last year
- 基于STM32F407开发板教程文档☆28Mar 7, 2020Updated 6 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- GraphRAG4OpenWebUI integrates Microsoft's GraphRAG technology into Open WebUI, providing a versatile information retrieval API. It combin…☆603Jan 10, 2025Updated last year
- 生成中文文字识别(OCR)的训练数据☆12Mar 2, 2020Updated 6 years ago
- ☆19Nov 9, 2024Updated last year
- Fork of turbo quant tom-tom and TurboQuant KV cache from domwox☆30Jul 7, 2026Updated 3 weeks ago
- M5Stack Cardputer(自带屏幕、键盘、麦克风、扬声器)上实现 MimiClaw☆49Mar 28, 2026Updated 4 months ago
- 基于vercel Serverless Functions搭建的无服务xss平台☆22Oct 29, 2024Updated last year
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆77Jul 14, 2026Updated 2 weeks ago
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆128Updated this week
- 自动化控制、可视化操作、自动连点器、流程自动化、RPA☆71Apr 5, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Proxmox Backup Server in Docker☆12Jan 17, 2026Updated 6 months ago
- 基于 PyQt5 的网络信息抓取工具,支持 Google/Bing/Baidu 多引擎搜索,自动抓取关键词相关内容并保存至本地 | A PyQt5-based web scraping tool that fetches keyword-related content fr…☆23May 22, 2026Updated 2 months ago
- rtmp/rtsp Qt Media☆21Jul 14, 2026Updated 2 weeks ago
- A lightweight llama.cpp launcher with GUI☆22Jun 12, 2026Updated last month
- LLM inference in C/C++☆2,203Jul 22, 2026Updated last week
- On cloud or premise, the AutoML controller for DeepCamera☆13Sep 1, 2022Updated 3 years ago
- People and asset tracking system. Hardware support : 1. Happy Bubble Presence Detectors 2.Raspberryp Pi 3. ESP32 kit with Wifi and Blueto…☆10Jun 6, 2018Updated 8 years ago
- ☆14Mar 7, 2023Updated 3 years ago
- 炫酷智慧城市效果(threejs加载geojson)☆10Mar 9, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- 使用小爱音箱播放音乐,音乐使用MusicFree插件在线搜索☆35May 31, 2026Updated last month
- Is it difficult to develop C++ high-concurrency server applications? Come and use XServer☆10Jun 13, 2024Updated 2 years ago
- Books management system based on python flask library.基于flask的图书管理系统支持电子书下载及纸质书的管理(借阅流转)☆12May 31, 2017Updated 9 years ago
- DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm☆19,409Updated this week
- web在线答题☆11Aug 14, 2018Updated 7 years ago
- Sherpa-onnx-tts-stt source for homeassisstant addon.☆26Nov 29, 2025Updated 8 months ago
- ☆20Oct 7, 2023Updated 2 years ago