Llama.cpp runner/swapper and proxy that emulates LMStudio / Ollama backends
☆60Aug 21, 2025Updated 11 months ago
Alternatives and similar repositories for llama-runner
Users that are interested in llama-runner are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- llama-swap + a minimal ollama compatible api☆61May 26, 2026Updated 2 months ago
- Simple node proxy for llama-server that enables MCP use☆19May 10, 2025Updated last year
- A simple, "Ollama-like" tool for managing and running GGUF language models from your terminal.☆25Jan 2, 2026Updated 7 months ago
- World's first Nintendo 3DS emulator for Apple devices based on Citra.☆18Apr 7, 2023Updated 3 years ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆148Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Testbench for llama.cpp llama-server☆15Aug 20, 2025Updated 11 months ago
- ☆24Dec 29, 2025Updated 7 months ago
- Personal voice assistant, with voice interruption and Twilio support☆18Feb 24, 2025Updated last year
- An open-source AI agent that lives in your terminal, privacy oriented, with no telemetry.☆77Updated this week
- Llama Server Launcher (llama.cpp/ik_llama) GUI☆122Jul 22, 2026Updated 3 weeks ago
- ☆58Oct 10, 2025Updated 10 months ago
- Local-first agent runtime for MCP workflows with explicit trust controls, replayable runs, and built-in evals.☆32Updated this week
- A TTS model capable of generating ultra-realistic dialogue in one pass.☆32May 1, 2025Updated last year
- ☆224May 7, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ACE-Step: A Step Towards Music Generation Foundation Model☆50May 20, 2025Updated last year
- Getting VibeVoice 7b working with 10 gb of vram.☆15Aug 31, 2025Updated 11 months ago
- Locally hosted AI Agent Python Tool To Generate Novel Research Hypothesis + Titles + Abstracts☆30Apr 30, 2025Updated last year
- llama.cpp fork with additional SOTA quants and improved performance☆22Updated this week
- 🔍📃 LLM-powered PDF Table Extractor☆19Jun 26, 2025Updated last year
- 🚀 FlexLLama - Lightweight self-hosted tool for running multiple llama.cpp server instances with OpenAI v1 API compatibility and multi-GP…☆60Jun 10, 2026Updated 2 months ago
- ☆16Oct 28, 2025Updated 9 months ago
- A python program that turns an LLM, running on Ollama, into an automated researcher, which will with a single query determine focus areas…☆13Nov 23, 2024Updated last year
- Generate Duolingo-style quiz courses from PDFs with spaced repetition, adaptive difficulty, and tutor chat.☆16Apr 6, 2026Updated 4 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- FlexAudioPrint is a Python-based app for transcribing audio to text using OpenAI's Whisper model. It offers a Gradio web interface and a …☆10Apr 22, 2026Updated 3 months ago
- A web application that converts speech to speech 100% private☆86Jun 3, 2025Updated last year
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆64Jul 7, 2026Updated last month
- ☆93Jul 7, 2025Updated last year
- ☆15Mar 18, 2026Updated 5 months ago
- The most feature-complete local AI workstation. Multi-GPU inference, integrated Stable Diffusion + ADetailer, voice cloning, research-gra…☆66Jul 28, 2026Updated 3 weeks ago
- Scripts and tools for optimizing quantizations in llama.cpp with GGUF imatrices.☆19Jan 10, 2025Updated last year
- An educational Rust project for exporting and running inference on Qwen3 LLM family☆43Aug 3, 2025Updated last year
- ☆19Jul 4, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- LexiCrawler is a powerful Go-based web crawling API meticulously designed to extract, clean, and transform web page content into a pristi…☆48Feb 27, 2025Updated last year
- German "Who Wants To Be A Millionaire" LLM Benchmarking.☆50Jul 24, 2026Updated 3 weeks ago
- ☆16Dec 16, 2024Updated last year
- Electron speech-to-speech app for your voice calls based on 100% locally run AI models☆35Jul 23, 2025Updated last year
- The easiest & fastest way to run LLMs in your home lab☆90Feb 23, 2026Updated 5 months ago
- Newman reporter allowing to decorate pull request with postman collection results.☆10Nov 2, 2023Updated 2 years ago
- Thireus's fork of llama.cpp with Cuda 12.8 and 13.3 release builds and Windows patch for loading more .gguf shards + llama-sweep-bench☆30Updated this week