☆119Apr 28, 2026Updated 4 months ago
Alternatives and similar repositories for qwen36-27b-single-3090
Users that are interested in qwen36-27b-single-3090 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Qwen3.6-27B on dual RTX 3090 — TP=2 recipe, vLLM nightly, MTP + fp8 KV, validated for concurrent serving☆58Apr 28, 2026Updated 4 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,182Updated this week
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆131Updated this week
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,832Updated this week
- Validated recipe for serving Qwen3.6-27B on a single RTX 5090 — full OpenAI API, vision, tool calling, MTP spec-decode☆73Apr 29, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Step-by-step guide: Qwen3.6 27B/35B on RTX PRO 6000 Blackwell — 120-200 tok/s with vLLM + MTP n=3☆22Aug 4, 2026Updated last month
- One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.☆228May 14, 2026Updated 3 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,027Updated this week
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆25Jan 30, 2026Updated 7 months ago
- Experimental implementation of DeepSeek v4 flaash in llama.cpp☆24Apr 30, 2026Updated 4 months ago
- Tool-calling quality benchmark for LLM serving stacks. 80+ deterministic scenarios testing multi-turn orchestration, safety boundaries, a…☆328Updated this week
- Experimental llama.cpp fork for inference research and development☆819Updated this week
- The mental model layer for agent-written code☆20Jan 21, 2026Updated 7 months ago
- Pure Go bash implementation for agent sandboxes☆47Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A modular framework for building massively parallel agentic systems☆32Sep 8, 2025Updated 11 months ago
- WinDbg Symbols Caching Proxy.☆18Updated this week
- Mixed-vendor GPU inference cluster manager with speculative decoding☆33Jul 2, 2026Updated 2 months ago
- Multi-GPU device selection for LTXV2 video generation in ComfyUI☆34Jan 10, 2026Updated 7 months ago
- NVIDIA Linux open GPU with P2P support☆478Updated this week
- A personal tutor extension for pi that adapts to your learning style, remembers what you're learning, and guides you with hints, projec…☆28Apr 19, 2026Updated 4 months ago
- Continuous Aggregation in Redis☆18Oct 23, 2023Updated 2 years ago
- A macOS app for running parallel AI agents in sandboxed local VMs☆21May 18, 2026Updated 3 months ago
- Agent-first Rust ASR orchestration stack: Bayesian backend routing across whisper.cpp/insanely-fast-whisper/whisper-diarization, real-tim…☆90Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A Windows driver used to facilitate DLL injection☆27Oct 29, 2017Updated 8 years ago
- ☆43May 4, 2026Updated 4 months ago
- In this repo, I developed a step-by-step pipeline for a standard MultiSpeaker Text-to-Speech system In general, I used Portaspeech as an…☆12Nov 24, 2023Updated 2 years ago
- A lightweight orchestration framework that piggybacks your local Agentic CLI setup. Intentionally simple. Yet powerful... like an army of…☆24Jun 21, 2026Updated 2 months ago
- macOS accessibility API showcase.☆11Jun 27, 2025Updated last year
- AI-powered dashboard builder — describe what you want in natural language and watch interactive charts, KPIs, and visualizations stream i…☆21Apr 2, 2026Updated 5 months ago
- common go code☆14Feb 22, 2021Updated 5 years ago
- The repository provides code for the paper RECE: Reduced Cross-Entropy Loss for Large-Catalogue Sequential Recommenders, CIKM'24☆11Oct 21, 2024Updated last year
- Add NextJS to Moleculer! 🎉☆11Jun 21, 2018Updated 8 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,292Updated this week
- Local benchmarking UI for LLMs and AI agents☆24Apr 13, 2026Updated 4 months ago
- Web research for your agents with smart and safe tooling + knowledge store☆34Updated this week
- An MCP-enabled Qwen3 0.6B demo with adjustable thinking budget, all in your browser!☆28Jun 2, 2025Updated last year
- Aggressive decode optimizations for Qwen3-0.6B on RTX 5090☆60Feb 25, 2026Updated 6 months ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,570Updated this week
- sanoTTS (सानो = 'small' in Nepali): a ~1.4M-param neural TTS that runs on a $3 chip or in the browser. Leads SCOREQ/UTMOS in the sub-15M …☆172Updated this week