vLLM fork for Tesla V100 (SM70) — extends 1CatAI's AWQ support and adds GGUF support
☆23Jun 20, 2026Updated 2 months ago
Alternatives and similar repositories for vllm-v100
Users that are interested in vllm-v100 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- V100 / SM70-focused vLLM engineering fork for modern LLM inference.☆1,010Updated this week
- Implementation of FlashAttention-2 for Nvidia Tesla V100 / Titan V☆206Jun 30, 2026Updated 2 months ago
- forked from vllm-project/flash-attention☆64May 9, 2026Updated 4 months ago
- A ComfyUI custom node enabling **Flash Attention 1** on legacy NVIDIA GPUs (Tesla V100, T4) that lack Compute Capability 8.0+ required by…☆25Feb 9, 2026Updated 7 months ago
- llama-cpp support turboquant and gemma 4☆21May 19, 2026Updated 3 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- TeamCopilot: Deploy AI agents for your team to automate business workflows and coding.☆16Jun 23, 2026Updated 2 months ago
- Unified KV-cache compression for LLM inference: 12 Python-native methods, guarded add-on composition and routing, analytical capacity sim…☆26Aug 22, 2026Updated 3 weeks ago
- Simple "Oscilloscope" on CH32V003J4M6☆16Dec 17, 2024Updated last year
- Superseded by github.com/poisonxa16/pxa — PXA, the set-and-forget engine for Pascal and Volta☆24Updated this week
- The htop for LLM inference see exactly where every GB of VRAM goes and get measured quantization savings.☆70Sep 1, 2026Updated 2 weeks ago
- Dynamic LLM model swapping system with Docker, vLLM integration, and GPU acceleration. Supports GGUF & Hugging Face models with automatic…☆23Mar 6, 2026Updated 6 months ago
- A professional ESP32-based DMR hotspot with MMDVM modem support, real-time web interface and BrandMeister network integration.☆24Apr 2, 2026Updated 5 months ago
- This project is specifically developed for V100, based on lmdeploy 0.12.1, and supports mainstream open-source models from Q4 2025 to Q1 …☆21Mar 18, 2026Updated 5 months ago
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆23Apr 3, 2026Updated 5 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- LuCI support for speedtest-web 内网测速网页版☆30Dec 7, 2021Updated 4 years ago
- various tools i've made for using mugen☆34Mar 16, 2017Updated 9 years ago
- Stop prompting your agents to behave. Start engineering them to.☆26Apr 17, 2026Updated 4 months ago
- Alternative IP Camera firmware from an open community☆19Jan 22, 2026Updated 7 months ago
- A fuzzer for ML compilers☆47Aug 7, 2026Updated last month
- An improved XLX reflector☆23Jul 30, 2025Updated last year
- ☆15Apr 16, 2026Updated 4 months ago
- A Python library to properly handle escaping of command line arguments in Windows' CMD.exe and Powershell.☆14Apr 26, 2023Updated 3 years ago
- My playground. Currently used as an auto builder for some custom board configs. EXPERIMENTAL CHANGES WARNING☆24Jan 2, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- The Billion Primes Project - Listing all prime numbers under 1 billion☆14Oct 5, 2020Updated 5 years ago
- Dual channel Arduino Nano milliwatt power meter for HF/VHF/UHF/SHF bands☆30Nov 1, 2024Updated last year
- A simple, observable code-writing agent builder in TypeScript.☆33Apr 9, 2025Updated last year
- llama.cpp fork with TurboQuant quantization (turbo2/3/4) and TriAttention GPU-accelerated KV cache pruning. 75 tok/s on Qwen3-8B / RTX 30…☆54Jul 2, 2026Updated 2 months ago
- A template for building ImmortalWrtwith GitHub Actions | 使用 GitHub Actions 云编译 ImmortalWrt☆11Dec 22, 2024Updated last year
- Low cost, low part count CH32V003-based oscilloscope☆29Dec 27, 2025Updated 8 months ago
- SX1255 as a software defined radio transceiver (repository for software)☆28Apr 2, 2026Updated 5 months ago
- M17 Analog Gateway by ESP32☆36Sep 19, 2023Updated 2 years ago
- Experimental collection of anycast routes☆18Apr 6, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Dshot for Raspberry Pi 5 using RP1 and piolib☆43Mar 22, 2025Updated last year
- Collaborative AI agent steering platform built on OpenCode☆16Mar 20, 2026Updated 5 months ago
- Adaptive Reasoning Engine for Efficient and Context-Aware Intelligence☆51Nov 15, 2025Updated 10 months ago
- ☆154Jun 13, 2026Updated 3 months ago
- 🤵 AIfred-Intelligence — self-hosted Multi-Agent Assistant with Debate Modes (Symposion/Tribunal), Voice (STT + Streaming-TTS), RAG with …☆37Updated this week
- ☆35Jan 30, 2025Updated last year
- 基于 FastAPI 的 DeepSeek Chat 反向代理,将 DeepSeek 网页版的 API 转换为 OpenAI 兼容格式。 支持流式/非流式对话、专家模式、深度思考(reasoning_content)、工具调用(DSML prompt injectio…☆141Updated this week