This repo contains the critical chat_template.jinja fix for Qwen3/3.5/3.6 to work smoothly in agentic task.
☆96May 2, 2026Updated 5 months ago
Alternatives and similar repositories for vLLM-Qwen3-3.5-3.6-chat-template-fix
Users that are interested in vLLM-Qwen3-3.5-3.6-chat-template-fix are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FlashQLA TileLang GDN kernels ported to NVIDIA Blackwell consumer (GB10 / DGX Spark)☆18Jun 5, 2026Updated 3 months ago
- Production-ready, bleeding-edge, upstream-first vLLM Docker image for the NVIDIA DGX Spark (GB10 / sm_121a), verified against real models…☆53Updated this week
- Compressed KV cache as cross-backend wire format for Metal + CUDA split inference over Thunderbolt 5☆16Apr 14, 2026Updated 5 months ago
- Local-first AI workflow orchestration for chaining models, agents, tools, and scripts into repeatable workflows.☆34Aug 10, 2026Updated last month
- Pi coding agent extensions with code intelligence (LSP + Tree-sitter AST), semantic refactoring, code review, web/Context7 docs, and ask-…☆96Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A database of knowledge around inference & training on GFX906 GPUs https://skyne98.github.io/wiki-gfx906/☆17Feb 21, 2026Updated 7 months ago
- Maic: A high-performance, MLX-optimized Local LLM server for Apple Silicon. OpenAI-compatible API for M1/M2/M3/M4 Macs and Mac Mini with …☆16Sep 14, 2026Updated 2 weeks ago
- ☆19Mar 19, 2026Updated 6 months ago
- ☆16Feb 3, 2026Updated 8 months ago
- Multimodal AI studio powered by Qwen3.6-35B-A3B. End-to-end web app exposing visual reasoning, image captioning, and document understandi…☆29Apr 23, 2026Updated 5 months ago
- ☆17May 21, 2026Updated 4 months ago
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆102Updated this week
- Your AI forgets everything between sessions. SAME fixes that. Local-first, no API keys, single binary.☆23Sep 3, 2026Updated 3 weeks ago
- Tool schemas get re-prefilled on every LLM request. ContextCache compiles them into a KV cache once, stores it on disk and reuses it acro…☆22Sep 16, 2026Updated 2 weeks ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,143Updated this week
- Qwen3.6-27B on dual RTX 3090 — TP=2 recipe, vLLM nightly, MTP + fp8 KV, validated for concurrent serving☆58Apr 28, 2026Updated 5 months ago
- Official Spark Arena Recipe Registry☆64Sep 3, 2026Updated last month
- ☆23Mar 31, 2025Updated last year
- A dedicated effort to make an optimized, bleeding edge vLLM image using Docker to support DGX comprehensively☆125Feb 22, 2026Updated 7 months ago
- Lightweight request runner☆18Dec 8, 2020Updated 5 years ago
- ☆202Sep 2, 2026Updated last month
- ☆17Aug 13, 2026Updated last month
- A collection of private CoreGraphics and SkyLight routines.☆11Feb 23, 2022Updated 4 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ABES Podcast Website☆10Aug 26, 2024Updated 2 years ago
- ☆28Dec 2, 2025Updated 10 months ago
- A transparent (O)llama proxy with model deployment aware routing which auto-manages multiple (O)llama instances in a given network.☆20Apr 14, 2026Updated 5 months ago
- A simpler self-hosted alternative to Open WebUI. Bring your own API keys or local models. Native Android and iOS clients.☆97Updated this week
- MCP Server for Namecheap☆20Jun 19, 2025Updated last year
- Docker configuration for running VLLM on dual DGX Sparks☆2,339Updated this week
- Automate your AI music video workflow with Synesthesia Engine. This local Gradio app bridges audio analysis, LLM-driven storytelling, and…☆54Aug 16, 2026Updated last month
- A GUI to manage 2025 ROG Flow Z13 system settings via z13ctl's API.☆37Aug 14, 2026Updated last month
- small script for managing google scholar alert emails☆11May 6, 2023Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A local dual-layer memory pattern for AI agents: a compact, human-readable markdown index paired with semantic retrieval from a local vec…☆62Updated this week
- Custom Olares Market source optimized for Olares One (RTX 5090M + Core Ultra 9 275HX)☆19Updated this week
- VellumForge2 is a Golang CLI for generating high-quality Direct Preference Optimization datasets via a hierarchical prompt pipeline with …☆22Feb 15, 2026Updated 7 months ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,805Updated this week
- A collection of intelligent AI agents built using Ollama, LangChain, and local LLMs — including chatbots, voice assistants, web scrapers,…☆20Jul 23, 2025Updated last year
- AEAD cipher based on ChaCha20 stream cipher and Poly1305 MAC☆10Feb 18, 2022Updated 4 years ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,333Updated this week