Qwen3.6-27B on dual RTX 3090 — TP=2 recipe, vLLM nightly, MTP + fp8 KV, validated for concurrent serving
☆57Apr 28, 2026Updated 3 months ago
Alternatives and similar repositories for qwen36-dual-3090
Users that are interested in qwen36-dual-3090 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆119Apr 28, 2026Updated 3 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,068Updated this week
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆43Jun 28, 2026Updated last month
- vLLM Docker Container for Qwen3.6 27b☆50Jun 15, 2026Updated 2 months ago
- Eclipse plugin for coding agents and other improvements☆17Jul 21, 2026Updated last month
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Pi extension that tracks bash tool token usage with live stats, grouping, and export☆23Feb 10, 2026Updated 6 months ago
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆132Updated this week
- One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.☆227May 14, 2026Updated 3 months ago
- Profile repo — categorized index of NVFP4 model releases, DGX Spark inference stacks, Apple Silicon MLX builds, the voice-AI stack, and t…☆39Jul 9, 2026Updated last month
- The mental model layer for agent-written code☆20Jan 21, 2026Updated 7 months ago
- Codemap generates a compact, token-aware map of a codebase: files, symbols, and markdown structure. Designed for feeding context to LLMs …☆40Jan 29, 2026Updated 6 months ago
- Connect Pi Agent Harness to Telegram☆17Mar 17, 2026Updated 5 months ago
- A Minimalistic Search Agent☆79May 12, 2026Updated 3 months ago
- LLM speculative inference server for consumer & heterogeneous hardware☆2,784Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A pi extension that suggests the user's likely next prompt.☆51Jun 2, 2026Updated 2 months ago
- ☆28Aug 8, 2026Updated 2 weeks ago
- Awesome Qiskit is a list of projects, tools, utilities, libraries and tutorials from a broad community of developers and researchers.☆31Jun 10, 2024Updated 2 years ago
- Generates breakthrough ideas from a single prompt through an 8 stage walkthrough, with optional research proposal paper.☆64Sep 30, 2025Updated 10 months ago
- NVIDIA Linux open GPU with P2P support☆441Updated this week
- A powerful MCP testing tool with multi-provider LLM support (Ollama, OpenAI, Claude, Gemini). Test, debug, and develop MCP servers with a…☆18Apr 28, 2026Updated 3 months ago
- Short Python script for parsing Defender VDM signature files.☆10Sep 22, 2024Updated last year
- Bleeding-edge ComfyUI for NVIDIA DGX Spark (GB10/Blackwell/sm_121a). CUDA 13 + SageAttention v3 (sm_121a) + NVFP4 + 14 custom-node packs …☆80Jun 28, 2026Updated last month
- A personal tutor extension for pi that adapts to your learning style, remembers what you're learning, and guides you with hints, projec…☆27Apr 19, 2026Updated 4 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Pi extension: a persistent second model that reviews the main agent's work each turn and injects concise advice inline.☆105Jul 11, 2026Updated last month
- ☆16Mar 21, 2025Updated last year
- Intellij plugin providing a simple browser embedded in tool window☆15Apr 5, 2026Updated 4 months ago
- Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.i…☆451Jul 3, 2026Updated last month
- Exposing the Neutrino EK: All the Naughty Bits (BSidesLV16)☆16Oct 10, 2016Updated 9 years ago
- Automatically rotate JPEG images on upload, according to Exif Orientation data for consistent, upright viewing on a web page.☆13May 19, 2016Updated 10 years ago
- A library for building advanced, optimised Ember.js apps☆25Jul 29, 2026Updated 3 weeks ago
- AI-powered dashboard builder — describe what you want in natural language and watch interactive charts, KPIs, and visualizations stream i…☆21Apr 2, 2026Updated 4 months ago
- Validated recipe for serving Qwen3.6-27B on a single RTX 5090 — full OpenAI API, vision, tool calling, MTP spec-decode☆73Apr 29, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Pi extension providing web-based terminal access via ghostty-web☆33Jul 13, 2026Updated last month
- SilverStripe module to allow users to "masquerade" as other users☆14May 4, 2026Updated 3 months ago
- RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink☆874Updated this week
- OpenCore EFI folder for running macOS Sonoma or newer on the Dell Optiplex 5050 Micro Small Form Factor PC.☆15Mar 18, 2026Updated 5 months ago
- Semantic search over past pi coding sessions — index, browse, and read your coding history☆22Jul 4, 2026Updated last month
- Configurable web access extension for pi that routes search, contents, answers, and research across Claude, Codex, Exa, Gemini, Parallel,…☆83Updated this week
- 4-5x faster Qwen3.5 on ASUS GX10 / DGX Spark — Hybrid INT4+FP8 + MTP via one shell script☆31Apr 16, 2026Updated 4 months ago