Qwen3.6-27B on dual RTX 3090 — TP=2 recipe, vLLM nightly, MTP + fp8 KV, validated for concurrent serving
☆58Apr 28, 2026Updated 5 months ago
Alternatives and similar repositories for qwen36-dual-3090
Users that are interested in qwen36-dual-3090 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆118Apr 28, 2026Updated 5 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,333Updated this week
- ☆15Apr 27, 2026Updated 5 months ago
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆44Jun 28, 2026Updated 3 months ago
- vLLM Docker Container for the latest Qwen 27b☆52Updated this week
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Eclipse plugin for coding agents and other improvements☆20Updated this week
- Pi extension that tracks bash tool token usage with live stats, grouping, and export☆23Feb 10, 2026Updated 7 months ago
- One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.☆230May 14, 2026Updated 4 months ago
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 4 months ago
- Profile repo — categorized index of NVFP4 model releases, DGX Spark inference stacks, Apple Silicon MLX builds, the voice-AI stack, and t…☆40Sep 21, 2026Updated last week
- The mental model layer for agent-written code☆20Jan 21, 2026Updated 8 months ago
- Codemap generates a compact, token-aware map of a codebase: files, symbols, and markdown structure. Designed for feeding context to LLMs …☆41Jan 29, 2026Updated 8 months ago
- Headless Matrix WebRTC voice AND video agent — auto-answers calls, bridges audio to any AI agent via PipeWire, optional camera-frame visi…☆20Jun 28, 2026Updated 3 months ago
- Harbor agent adapter for pi coding agent to run Terminal-Bench evaluations☆30Dec 1, 2025Updated 10 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,887Updated this week
- A pi extension that suggests the user's likely next prompt.☆52Jun 2, 2026Updated 4 months ago
- Yet Another (LLM) Web UI, made with Gemini☆12Dec 25, 2024Updated last year
- ☆29Aug 8, 2026Updated last month
- Test Environment Booking tool☆14Nov 16, 2020Updated 5 years ago
- Awesome Qiskit is a list of projects, tools, utilities, libraries and tutorials from a broad community of developers and researchers.☆31Jun 10, 2024Updated 2 years ago
- ☆20Oct 23, 2025Updated 11 months ago
- Short Python script for parsing Defender VDM signature files.☆10Sep 22, 2024Updated 2 years ago
- Bleeding-edge ComfyUI for NVIDIA DGX Spark (GB10/Blackwell/sm_121a). CUDA 13 + SageAttention v3 (sm_121a) + NVFP4 + 14 custom-node packs …☆86Jun 28, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A personal tutor extension for pi that adapts to your learning style, remembers what you're learning, and guides you with hints, projec…☆29Updated this week
- Pi extension: a persistent second model that reviews the main agent's work each turn and injects concise advice inline.☆111Aug 24, 2026Updated last month
- LLM frontend for roleplay☆80Sep 26, 2026Updated last week
- Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.i…☆469Sep 15, 2026Updated 2 weeks ago
- Headless voice interface for the Pi Coding Agent☆81Feb 13, 2026Updated 7 months ago
- Exposing the Neutrino EK: All the Naughty Bits (BSidesLV16)☆16Oct 10, 2016Updated 9 years ago
- Self-improving development workflows for pi coding agent. Subagent orchestration, TDD, systematic debugging. Adapted from obra/superpower…☆36Mar 30, 2026Updated 6 months ago
- ☆17Feb 15, 2026Updated 7 months ago
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆1,117Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Semantic search over past pi coding sessions — index, browse, and read your coding history☆23Sep 24, 2026Updated last week
- An MCP-enabled Qwen3 0.6B demo with adjustable thinking budget, all in your browser!☆28Jun 2, 2025Updated last year
- The web, from your terminal. Search, extract pages, get grounded answers, and run deep research with your choice of providers. CLI, TypeS…☆86Updated this week
- 4-5x faster Qwen3.5 on ASUS GX10 / DGX Spark — Hybrid INT4+FP8 + MTP via one shell script☆32Apr 16, 2026Updated 5 months ago
- Stable Diffusion in pure C/C++☆17Jun 21, 2026Updated 3 months ago
- Zero-instrumentation LLM API and MCP tracer for your agents powered by eBPF — latency, tokens, and tool use in realtime☆18Mar 16, 2026Updated 6 months ago
- pi extension to search the Web search via your Perplexity Pro/Max subscription☆31Updated this week