Llama.cpp runner/swapper and proxy that emulates LMStudio / Ollama backends
☆61Aug 21, 2025Updated last year
Alternatives and similar repositories for llama-runner
Users that are interested in llama-runner are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Simple node proxy for llama-server that enables MCP use☆19May 10, 2025Updated last year
- llama-swap + a minimal ollama compatible api☆62Updated this week
- A simple, "Ollama-like" tool for managing and running GGUF language models from your terminal.☆25Jan 2, 2026Updated 8 months ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆159Aug 21, 2026Updated 2 weeks ago
- Testbench for llama.cpp llama-server☆15Aug 20, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆24Dec 29, 2025Updated 8 months ago
- Personal voice assistant, with voice interruption and Twilio support☆18Feb 24, 2025Updated last year
- An open-source AI agent that lives in your terminal, privacy oriented, with no telemetry.☆79Updated this week
- Llama Server Launcher (llama.cpp/ik_llama) GUI☆125Jul 22, 2026Updated last month
- ☆59Oct 10, 2025Updated 10 months ago
- Local-first agent runtime for MCP workflows with explicit trust controls, replayable runs, and built-in evals.☆32Aug 15, 2026Updated 3 weeks ago
- A TTS model capable of generating ultra-realistic dialogue in one pass.☆32May 1, 2025Updated last year
- ☆226May 7, 2025Updated last year
- ACE-Step: A Step Towards Music Generation Foundation Model☆50May 20, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Getting VibeVoice 7b working with 10 gb of vram.☆15Aug 31, 2025Updated last year
- Locally hosted AI Agent Python Tool To Generate Novel Research Hypothesis + Titles + Abstracts☆30Apr 30, 2025Updated last year
- llama.cpp fork with additional SOTA quants and improved performance☆22Aug 31, 2026Updated last week
- 🔍📃 LLM-powered PDF Table Extractor☆19Jun 26, 2025Updated last year
- 🚀 FlexLLama - Lightweight self-hosted tool for running multiple llama.cpp server instances with OpenAI v1 API compatibility and multi-GP…☆59Jun 10, 2026Updated 2 months ago
- ☆16Oct 28, 2025Updated 10 months ago
- A python program that turns an LLM, running on Ollama, into an automated researcher, which will with a single query determine focus areas…☆13Nov 23, 2024Updated last year
- Generate Duolingo-style quiz courses from PDFs with spaced repetition, adaptive difficulty, and tutor chat.☆16Apr 6, 2026Updated 5 months ago
- FlexAudioPrint is a Python-based app for transcribing audio to text using OpenAI's Whisper model. It offers a Gradio web interface and a …☆10Apr 22, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A web application that converts speech to speech 100% private☆87Jun 3, 2025Updated last year
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆67Updated this week
- ☆93Jul 7, 2025Updated last year
- Natural language control for Python CLI tools using locally-trained SLMs (CPU inference)☆33Aug 15, 2026Updated 3 weeks ago
- ☆15Mar 18, 2026Updated 5 months ago
- Anthropic's Contextual Retrieval implementation with visual chunk comparison. Preview context enrichment before/after embedding.☆31Sep 25, 2025Updated 11 months ago
- Scripts and tools for optimizing quantizations in llama.cpp with GGUF imatrices.☆19Jan 10, 2025Updated last year
- An educational Rust project for exporting and running inference on Qwen3 LLM family☆44Aug 3, 2025Updated last year
- LexiCrawler is a powerful Go-based web crawling API meticulously designed to extract, clean, and transform web page content into a pristi…☆48Feb 27, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- German "Who Wants To Be A Millionaire" LLM Benchmarking.☆50Updated this week
- ☆16Dec 16, 2024Updated last year
- A collection of high quality huggingface datasets.☆30Apr 19, 2026Updated 4 months ago
- Electron speech-to-speech app for your voice calls based on 100% locally run AI models☆35Jul 23, 2025Updated last year
- The easiest & fastest way to run LLMs in your home lab☆90Feb 23, 2026Updated 6 months ago
- Thireus's fork of llama.cpp with Cuda 12.8 and 13.3 release builds and Windows patch for loading more .gguf shards + llama-sweep-bench☆30Updated this week
- ☆214Sep 7, 2025Updated last year