☆117Apr 28, 2026Updated 2 months ago
Alternatives and similar repositories for qwen36-27b-single-3090
Users that are interested in qwen36-27b-single-3090 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Qwen3.6-27B on dual RTX 3090 — TP=2 recipe, vLLM nightly, MTP + fp8 KV, validated for concurrent serving☆56Apr 28, 2026Updated 2 months ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆1,778Updated this week
- vLLM Docker Container for Qwen3.6 27b☆51Jun 15, 2026Updated last month
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆125Updated this week
- Fast LLM speculative inference server for consumer hardware.☆2,673Updated this week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Validated recipe for serving Qwen3.6-27B on a single RTX 5090 — full OpenAI API, vision, tool calling, MTP spec-decode☆68Apr 29, 2026Updated 2 months ago
- One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.☆224May 14, 2026Updated 2 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆801Updated this week
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆24Jan 30, 2026Updated 5 months ago
- Connect Chromium to LLM agents via token-efficient DOM compression. 50-200 tokens per page.☆20Jul 17, 2026Updated last week
- Experimental implementation of DeepSeek v4 flaash in llama.cpp☆23Apr 30, 2026Updated 2 months ago
- LLAMA Turboquant implementation with CUDA support☆710Updated this week
- Tool-calling quality benchmark for LLM serving stacks. 80+ deterministic scenarios testing multi-turn orchestration, safety boundaries, a…☆247Updated this week
- Pure Go bash implementation for agent sandboxes☆40Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- WinDbg Symbols Caching Proxy.☆18Updated this week
- Plugin for ida pro that copies RVA under cursor to clipboard.☆18Jul 28, 2023Updated 2 years ago
- Multi-GPU device selection for LTXV2 video generation in ComfyUI☆32Jan 10, 2026Updated 6 months ago
- NVIDIA Linux open GPU with P2P support☆355Jul 13, 2026Updated last week
- A personal tutor extension for pi that adapts to your learning style, remembers what you're learning, and guides you with hints, projec…☆27Apr 19, 2026Updated 3 months ago
- 72 AI Agent Skills for Manus - Properly formatted for direct import☆19Mar 10, 2026Updated 4 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆2,960Updated this week
- collection of 2.4 mods i've made☆10Dec 5, 2022Updated 3 years ago
- CoAct-1: Computer-using Agents with Coding as Actions☆27Jun 2, 2026Updated last month
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Chatons is a desktop AI workspace for coding and project workflows: it lets you chat with multiple AI providers, pick scoped or full mod…☆18May 27, 2026Updated last month
- ☆44May 4, 2026Updated 2 months ago
- C++ version of pyannote audio overlapped speech detection pipeline☆13Feb 14, 2024Updated 2 years ago
- An environment manager for ComfyUI☆18Jun 28, 2026Updated 3 weeks ago
- The code used for the documentation website.☆19Jul 15, 2026Updated last week
- Wiring distribution box for multiple toolheads on Voron2☆11Sep 16, 2023Updated 2 years ago
- Randomizer for Quest for Glory 1 EGA☆16Nov 12, 2023Updated 2 years ago
- ☆17Feb 15, 2026Updated 5 months ago
- Add NextJS to Moleculer! 🎉☆11Jun 21, 2018Updated 8 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Fused TBQ4 Flash Attention + MTP + Shared Tensors for llama.cpp — 82+ tok/s with lossless 4.25 bpv KV cache at 200K context on RTX 4090☆92May 17, 2026Updated 2 months ago
- Local benchmarking UI for LLMs and AI agents☆20Apr 13, 2026Updated 3 months ago
- Eurostat Big Data Hackathon 2021☆15Jan 15, 2026Updated 6 months ago
- Deploy Apollo HF space locally☆40Dec 16, 2024Updated last year
- An MCP-enabled Qwen3 0.6B demo with adjustable thinking budget, all in your browser!☆28Jun 2, 2025Updated last year
- Aggressive decode optimizations for Qwen3-0.6B on RTX 5090☆54Feb 25, 2026Updated 5 months ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,129Updated this week