Llama.cpp runner/swapper and proxy that emulates LMStudio / Ollama backends
☆60Aug 21, 2025Updated 11 months ago
Alternatives and similar repositories for llama-runner
Users that are interested in llama-runner are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- llama-swap + a minimal ollama compatible api☆60May 26, 2026Updated 2 months ago
- Simple node proxy for llama-server that enables MCP use☆19May 10, 2025Updated last year
- A simple, "Ollama-like" tool for managing and running GGUF language models from your terminal.☆25Jan 2, 2026Updated 6 months ago
- Produce your own Dynamic 3.0 Quants and achieve optimum accuracy & SOTA quantization performance! Input a target size and the toolchain w…☆146Updated this week
- Testbench for llama.cpp llama-server☆15Aug 20, 2025Updated 11 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆24Dec 29, 2025Updated 6 months ago
- Personal voice assistant, with voice interruption and Twilio support☆18Feb 24, 2025Updated last year
- An open-source AI agent that lives in your terminal, privacy oriented, with no telemetry.☆74Updated this week
- Llama Server Launcher (llama.cpp/ik_llama) GUI☆123Updated this week
- ☆58Oct 10, 2025Updated 9 months ago
- Local-first agent runtime for MCP workflows with explicit trust controls, replayable runs, and built-in evals.☆32Jul 4, 2026Updated 3 weeks ago
- A TTS model capable of generating ultra-realistic dialogue in one pass.☆32May 1, 2025Updated last year
- ☆224May 7, 2025Updated last year
- ACE-Step: A Step Towards Music Generation Foundation Model☆50May 20, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Getting VibeVoice 7b working with 10 gb of vram.☆15Aug 31, 2025Updated 10 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆22Jul 10, 2026Updated 2 weeks ago
- 🔍📃 LLM-powered PDF Table Extractor☆19Jun 26, 2025Updated last year
- 🚀 FlexLLama - Lightweight self-hosted tool for running multiple llama.cpp server instances with OpenAI v1 API compatibility and multi-GP…☆59Jun 10, 2026Updated last month
- ☆16Oct 28, 2025Updated 9 months ago
- Locally hosted AI Agent Python Tool To Generate Novel Research Hypothesis + Titles + Abstracts☆30Apr 30, 2025Updated last year
- A python program that turns an LLM, running on Ollama, into an automated researcher, which will with a single query determine focus areas…☆13Nov 23, 2024Updated last year
- Generate Duolingo-style quiz courses from PDFs with spaced repetition, adaptive difficulty, and tutor chat.☆16Apr 6, 2026Updated 3 months ago
- FlexAudioPrint is a Python-based app for transcribing audio to text using OpenAI's Whisper model. It offers a Gradio web interface and a …☆10Apr 22, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A web application that converts speech to speech 100% private☆86Jun 3, 2025Updated last year
- Profile-guided GPU kernel optimizer for AMD/RDNA3. Auto-tunes llama.cpp MMVQ kernels per model shape. 2x decode speedup on 7900 XTX.☆64Jul 7, 2026Updated 3 weeks ago
- ☆93Jul 7, 2025Updated last year
- Natural language control for Python CLI tools using locally-trained SLMs (CPU inference)☆32Apr 10, 2026Updated 3 months ago
- ☆15Mar 18, 2026Updated 4 months ago
- Anthropic's Contextual Retrieval implementation with visual chunk comparison. Preview context enrichment before/after embedding.☆30Sep 25, 2025Updated 10 months ago
- The most feature-complete local AI workstation. Multi-GPU inference, integrated Stable Diffusion + ADetailer, voice cloning, research-gra…☆64Updated this week
- Scripts and tools for optimizing quantizations in llama.cpp with GGUF imatrices.☆19Jan 10, 2025Updated last year
- An educational Rust project for exporting and running inference on Qwen3 LLM family☆44Aug 3, 2025Updated 11 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆19Jul 4, 2025Updated last year
- axseem's Linux Workstation Configuration | Mirror of https://codeberg.org/axseem/dots☆20Updated this week
- ☆16Dec 16, 2024Updated last year
- A collection of high quality huggingface datasets.☆29Apr 19, 2026Updated 3 months ago
- Electron speech-to-speech app for your voice calls based on 100% locally run AI models☆35Jul 23, 2025Updated last year
- The easiest & fastest way to run LLMs in your home lab☆90Feb 23, 2026Updated 5 months ago
- ☆214Sep 7, 2025Updated 10 months ago