Linux & Powershell scripts to easily set up and run the Qwen 3.5 series locally on Windows and Linux with llama.cpp.
☆92Apr 28, 2026Updated 3 months ago
Alternatives and similar repositories for local-qwen3-coder-env
Users that are interested in local-qwen3-coder-env are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆165Updated this week
- Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with multi-GPU tensor-split, MoE expert placem…☆260Updated this week
- ☆44May 4, 2026Updated 2 months ago
- a fast and lightweight distributed background task processing framework with seamless scheduling.☆15Mar 30, 2026Updated 3 months ago
- this is an easy way to make ai podcast useing ai loccaly like ollama and the tts of piper☆16Feb 17, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A fully local document intelligence system that allows users to build a persistent private knowledge base from documents and query it usi…☆16Apr 15, 2026Updated 3 months ago
- Practical local LLM recipes and benchmarks for RTX 5060 Ti setups☆91Jul 8, 2026Updated 3 weeks ago
- llama.cpp fork with additional SOTA quants and improved performance☆2,969Updated this week
- Messy repo filled with messy tests about hardware and LLMs. Built for me, public for you.☆43Jun 1, 2026Updated last month
- Thireus's fork of llama.cpp with Cuda 12.8 and 13.3 release builds and Windows patch for loading more .gguf shards + llama-sweep-bench☆30Updated this week
- llama-swap + a minimal ollama compatible api☆60May 26, 2026Updated 2 months ago
- Llama.cpp runner/swapper and proxy that emulates LMStudio / Ollama backends☆60Aug 21, 2025Updated 11 months ago
- Smart home assistant powered by an SLM☆17May 16, 2026Updated 2 months ago
- Copilot with deepseek and more...☆13Mar 7, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A harness optimized to smaller LLMs☆2,061Updated this week
- ☆36Aug 21, 2025Updated 11 months ago
- Smart OpenAI‑compatible proxy for llama.cpp: manages slots, saves/restores KV cache to disk, routes requests by prefix similarity, and pr…☆49Nov 14, 2025Updated 8 months ago
- Custom Template for checking the availiability of an entity.☆13Oct 31, 2025Updated 8 months ago
- Kon is a minimal coding agent (and also a highly opinionated one)☆340Updated this week
- An electron Wrapper for Open-Interpreter for the lablab.ai hackathon☆12Oct 14, 2023Updated 2 years ago
- An open-source AI agent that lives in your terminal, privacy oriented, with no telemetry.☆74Updated this week
- Scripts and tools for optimizing quantizations in llama.cpp with GGUF imatrices.☆19Jan 10, 2025Updated last year
- Compare tables within or across databases☆16Jun 4, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- GPU-accelerated voice assistant — local LLM, fine-tuned Whisper, Kokoro TTS, AMD ROCm☆17Apr 2, 2026Updated 3 months ago
- ☆16Dec 16, 2024Updated last year
- Deploying full-stack on-prem deep research agent that can be run entirely on a local machine for $0!☆34Nov 8, 2025Updated 8 months ago
- An OpenAI-compatible ASR/STT API server powered by Meta's omnilingual-asr model. Supports real-time streaming via WebSocket and batch tra…☆18Jan 2, 2026Updated 6 months ago
- One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.☆222May 14, 2026Updated 2 months ago
- LLM inference in C/C++☆116Updated this week
- ☆19Updated this week
- ☆20Jul 4, 2025Updated last year
- ☆12Aug 3, 2024Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Huntress API☆11May 26, 2022Updated 4 years ago
- ☆20Mar 17, 2026Updated 4 months ago
- ☆12Sep 9, 2024Updated last year
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆815Updated this week
- ROCm/AMD GPU benchmark suite for llama.cpp, whisper.cpp, PyTorch☆18Feb 4, 2026Updated 5 months ago
- ☆14Oct 9, 2023Updated 2 years ago
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 2 months ago