Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with multi-GPU tensor-split, MoE expert placement, measured flag tuning (AI Tune), hardware-matched HuggingFace downloads, and crash recovery. An Ollama alternative for multi-GPU rigs.
☆255Jul 19, 2026Updated this week
Alternatives and similar repositories for ggrun
Users that are interested in ggrun are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆789Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆2,943Updated this week
- Linux & Powershell scripts to easily set up and run the Qwen 3.5 series locally on Windows and Linux with llama.cpp.☆90Apr 28, 2026Updated 2 months ago
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆125Updated this week
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆1,751Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,066Updated this week
- Proxy for OpenAI☆16Sep 2, 2025Updated 10 months ago
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,054Updated this week
- vLLM Docker Container for Qwen3.6 27b☆51Jun 15, 2026Updated last month
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 2 months ago
- Fast LLM speculative inference server for consumer hardware.☆2,668Updated this week
- Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware☆483Updated this week
- Llama.cpp launcher with integrated huggingface☆59Updated this week
- LLM inference in C/C++☆83Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆164Updated this week
- Human-AI Document Standard — lightweight convention for AI-optimized technical documentation☆28Jul 9, 2026Updated last week
- Run autoresearch on any NVIDIA GPUs (Works on 2-4GB+ Cards)☆25Mar 22, 2026Updated 3 months ago
- ☆44May 4, 2026Updated 2 months ago
- A harness optimized to smaller LLMs☆1,778Updated this week
- ☆53Updated this week
- SmarterRouter: An intelligent LLM gateway and VRAM-aware router for Ollama, llama.cpp, and OpenAI. Features semantic caching, model profi…☆145May 10, 2026Updated 2 months ago
- Fulloch - The Fully Local Home Voice Assistant☆121Updated this week
- A lightweight, prompt-driven MCP web research server for high-quality LLM powered information extraction.☆138Apr 10, 2026Updated 3 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆146Updated this week
- A TUI around llama.cpp for running, managing, and benchmarking local GGUF models and launching the pi coding agent against your local ser…☆24Updated this week
- RDNA-native LLM inference engine in Rust.☆484Updated this week
- Coding agent with cool features.☆16Jun 6, 2026Updated last month
- the composable multi-agent shell☆447Updated this week
- Open-source framework for superagents.☆91Jun 16, 2026Updated last month
- Llama.cpp runner/swapper and proxy that emulates LMStudio / Ollama backends☆60Aug 21, 2025Updated 11 months ago
- Local-first AI workflow orchestration for chaining models, agents, tools, and scripts into repeatable workflows.☆28Jul 14, 2026Updated last week
- Generate Duolingo-style quiz courses from PDFs with spaced repetition, adaptive difficulty, and tutor chat.☆16Apr 6, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, a…☆798Updated this week
- CTX - Context Runtime Engine for Coding Agents☆139May 11, 2026Updated 2 months ago
- Kon is a minimal coding agent (and also a highly opinionated one)☆337Updated this week
- Your AI's anchor to reality. ⚓☆60Jul 9, 2026Updated last week
- ☆44Apr 26, 2026Updated 2 months ago
- ☆16Mar 15, 2026Updated 4 months ago
- Using LLMs for iteratively exploring the solution search space at scale.☆732Jul 13, 2026Updated last week