☆119Apr 28, 2026Updated 3 months ago
Alternatives and similar repositories for qwen36-27b-single-3090
Users that are interested in qwen36-27b-single-3090 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Qwen3.6-27B on dual RTX 3090 — TP=2 recipe, vLLM nightly, MTP + fp8 KV, validated for concurrent serving☆57Apr 28, 2026Updated 3 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆1,932Updated this week
- LLM speculative inference server for consumer & heterogeneous hardware☆2,739Updated this week
- Validated recipe for serving Qwen3.6-27B on a single RTX 5090 — full OpenAI API, vision, tool calling, MTP spec-decode☆72Apr 29, 2026Updated 3 months ago
- Step-by-step guide: Qwen3.6 27B/35B on RTX PRO 6000 Blackwell — 120-200 tok/s with vLLM + MTP n=3☆19Aug 4, 2026Updated last week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Reproducible llama.cpp configs + per-category quality benches for Qwen3.6-27B on a single RTX 4090. Winners, dead ends, and the silent-co…☆23Apr 26, 2026Updated 3 months ago
- One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.☆227May 14, 2026Updated 2 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆860Updated this week
- Connect Chromium to LLM agents via token-efficient DOM compression. 50-200 tokens per page.☆21Updated this week
- Experimental implementation of DeepSeek v4 flaash in llama.cpp☆24Apr 30, 2026Updated 3 months ago
- Experimental llama.cpp fork for inference research and development☆740Updated this week
- Tool-calling quality benchmark for LLM serving stacks. 80+ deterministic scenarios testing multi-turn orchestration, safety boundaries, a…☆288Updated this week
- The mental model layer for agent-written code☆20Jan 21, 2026Updated 6 months ago
- Plugin for ida pro that copies RVA under cursor to clipboard.☆18Jul 28, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Mixed-vendor GPU inference cluster manager with speculative decoding☆31Jul 2, 2026Updated last month
- Exercise files for the book titled Functional Programming in PHP☆10Jun 11, 2022Updated 4 years ago
- ☆17Jul 16, 2026Updated 3 weeks ago
- Multi-GPU device selection for LTXV2 video generation in ComfyUI☆34Jan 10, 2026Updated 7 months ago
- ATTINY85 High-Voltage Flash Programmer☆14May 27, 2020Updated 6 years ago
- A puppeteer-extra plugin to solve Amazon captchas using Tessaract.JS.☆15May 16, 2024Updated 2 years ago
- Event Engine Docs☆13Jan 2, 2022Updated 4 years ago
- Script to make an automatic purchase on Amazon☆10Nov 30, 2020Updated 5 years ago
- Zip Prompt Language - compress heavy system prompts by ≥ 80 % token reduction☆15Jun 6, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- NVIDIA Linux open GPU with P2P support☆396Updated this week
- ValueObjects toolkit, for making your own value objects quickly and easily☆10Nov 10, 2017Updated 8 years ago
- Bash scripts to backup and restore proxmox server.☆13Aug 7, 2025Updated last year
- A novel approach for transformer model introspection that enables saving, compressing, and manipulating internal thought states for advan…☆34Mar 22, 2026Updated 4 months ago
- Amazon Captcha solver in pure Rust☆17Nov 16, 2023Updated 2 years ago
- High-Performance Proxy Server with Automatic IPv6 Rotation☆41Feb 18, 2026Updated 5 months ago
- Agent-first Rust ASR orchestration stack: Bayesian backend routing across whisper.cpp/insanely-fast-whisper/whisper-diarization, real-tim…☆46Updated this week
- A macOS app for running parallel AI agents in sandboxed local VMs☆20May 18, 2026Updated 2 months ago
- Local LLM inference on RTX 4050 6GB VRAM using TurboQuant llama.cpp with Qwen3.6 35B A3B GGUF☆24Jul 10, 2026Updated last month
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- llama.cpp fork with additional SOTA quants and improved performance☆3,028Updated this week
- Chatons is a desktop AI workspace for coding and project workflows: it lets you chat with multiple AI providers, pick scoped or full mod…☆18May 27, 2026Updated 2 months ago
- ☆44May 4, 2026Updated 3 months ago
- In this repo, I developed a step-by-step pipeline for a standard MultiSpeaker Text-to-Speech system In general, I used Portaspeech as an…☆12Nov 24, 2023Updated 2 years ago
- A collections package for various use cases (supports strict typing)☆18Jan 23, 2026Updated 6 months ago
- A lightweight orchestration framework that piggybacks your local Agentic CLI setup. Intentionally simple. Yet powerful... like an army of…☆23Jun 21, 2026Updated last month
- Opinionated Elasticsearch query language for Rust☆11Jul 24, 2021Updated 5 years ago