llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.
☆275Sep 19, 2026Updated this week
Alternatives and similar repositories for ggrun
Users that are interested in ggrun are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,105Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆3,251Updated this week
- Linux & Powershell scripts to easily set up and run the Qwen 3.5 series locally on Windows and Linux with llama.cpp.☆111Aug 22, 2026Updated 3 weeks ago
- Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currentl…☆2,278Updated this week
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆132Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs☆1,472Updated this week
- Proxy for OpenAI☆16Sep 2, 2025Updated last year
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,709Updated this week
- Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware☆521Updated this week
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,870Updated this week
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 4 months ago
- Llama.cpp launcher with integrated huggingface☆72Sep 8, 2026Updated last week
- Human-AI Document Standard — lightweight convention for AI-optimized technical documentation☆28Jul 9, 2026Updated 2 months ago
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆177Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Run autoresearch on any NVIDIA GPUs (Works on 2-4GB+ Cards)☆25Mar 22, 2026Updated 5 months ago
- ☆43May 4, 2026Updated 4 months ago
- ☆79Updated this week
- A harness optimized to smaller LLMs☆2,608Updated this week
- SmarterRouter: An intelligent LLM gateway and VRAM-aware router for Ollama, llama.cpp, and OpenAI. Features semantic caching, model profi…☆152May 10, 2026Updated 4 months ago
- AI-assisted software development methodology for solo builders☆30Updated this week
- Fulloch - The Fully Local Home Voice Assistant☆154Updated this week
- Full AI context and content layer for coding agents over one MCP server — tree-sitter code-map, document RAG, shared memory, multi-agent …☆102Updated this week
- A lightweight, prompt-driven MCP web research server for high-quality LLM powered information extraction.☆139Apr 10, 2026Updated 5 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A TUI around llama.cpp for running, managing, and benchmarking local GGUF models and launching the pi coding agent against your local ser…☆31Updated this week
- RDNA-native LLM inference engine in Rust.☆635Updated this week
- Coding agent with cool features.☆19Aug 23, 2026Updated 3 weeks ago
- the composable multi-agent shell☆482Sep 10, 2026Updated last week
- Llama.cpp runner/swapper and proxy that emulates LMStudio / Ollama backends☆61Aug 21, 2025Updated last year
- Local-first AI workflow orchestration for chaining models, agents, tools, and scripts into repeatable workflows.☆33Aug 10, 2026Updated last month
- Open-source framework for superagents.☆113Sep 11, 2026Updated last week
- Generate Duolingo-style quiz courses from PDFs with spaced repetition, adaptive difficulty, and tutor chat.☆18Apr 6, 2026Updated 5 months ago
- CTX - Context Runtime Engine for Coding Agents☆143May 11, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, a…☆2,907Updated this week
- Your AI's anchor to reality. ⚓☆61Updated this week
- LLM inference in C/C++☆2,395Updated this week
- Using LLMs for iteratively exploring the solution search space at scale.☆739Sep 5, 2026Updated 2 weeks ago
- Reproducible llama.cpp configs + per-category quality benches for Qwen3.6-27B on a single RTX 4090. Winners, dead ends, and the silent-co…☆24Apr 26, 2026Updated 4 months ago
- ☆213Jan 5, 2026Updated 8 months ago
- TRELLIS.2 image-to-3D in C++/GGML (CUDA + Vulkan), with a resident HTTP server☆312Updated this week