llama-swap + a minimal ollama compatible api
☆60May 26, 2026Updated last month
Alternatives and similar repositories for llama-swappo
Users that are interested in llama-swappo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Llama.cpp runner/swapper and proxy that emulates LMStudio / Ollama backends☆60Aug 21, 2025Updated 11 months ago
- ☆99Mar 28, 2026Updated 3 months ago
- ☆57Oct 10, 2025Updated 9 months ago
- A proxy that hosts multiple single-model runners such as LLama.cpp and vLLM☆12May 30, 2025Updated last year
- A simple, "Ollama-like" tool for managing and running GGUF language models from your terminal.☆25Jan 2, 2026Updated 6 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Chat WebUI is an easy-to-use user interface for interacting with AI, and it comes with multiple useful built-in tools such as web search …☆52Feb 10, 2026Updated 5 months ago
- An fully autonomous agent that accesses the browser and performs tasks.☆18Apr 25, 2025Updated last year
- empirically chooses -ngl param for llama.cpp☆20Mar 19, 2025Updated last year
- Yet Another (LLM) Web UI, made with Gemini☆12Dec 25, 2024Updated last year
- ☆18Jul 1, 2025Updated last year
- ☆20Jul 4, 2025Updated last year
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,066Updated this week
- Run Orpheus 3B Locally with Gradio UI, Standalone App☆24Apr 1, 2025Updated last year
- ☆93Jul 7, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- VellumForge2 is a Golang CLI for generating high-quality Direct Preference Optimization datasets via a hierarchical prompt pipeline with …☆20Feb 15, 2026Updated 5 months ago
- "Pacha" TUI (Text User Interface) is a JavaScript application that utilizes the "blessed" library. It serves as a frontend for llama.cpp …☆38Aug 3, 2023Updated 2 years ago
- Bookmarklet to pull and run hugging face GGUF models in Ollama☆18Oct 17, 2024Updated last year
- A custom LiteLLM provider enabling local execution of Hugging Face models with streaming, quantization, and async support☆30Jun 22, 2025Updated last year
- A command-line client for Bluesky☆17May 29, 2026Updated last month
- ☆20Sep 28, 2024Updated last year
- A daemon that automatically manages the performance states of NVIDIA GPUs.☆136Feb 24, 2026Updated 4 months ago
- Running Microsoft's BitNet inference framework via FastAPI, Uvicorn and Docker.☆39Jul 2, 2025Updated last year
- CompChomper is a framework for measuring how LLMs perform at code completion.☆21Apr 29, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆21Jan 25, 2025Updated last year
- vLLM Docker Container for Qwen3.6 27b☆51Jun 15, 2026Updated last month
- 🚀 FlexLLama - Lightweight self-hosted tool for running multiple llama.cpp server instances with OpenAI v1 API compatibility and multi-GP…☆59Jun 10, 2026Updated last month
- Enable tool/function calling for any LLM, in OpenAI and Ollama API formats, adding universal function calling to models without native su…☆77Dec 9, 2025Updated 7 months ago
- Electron speech-to-speech app for your voice calls based on 100% locally run AI models☆35Jul 23, 2025Updated 11 months ago
- Fast LLM swapping with sleep/wake support, compatible with vllm, llama.cpp, etc. llama-swap fork.☆51Apr 5, 2026Updated 3 months ago
- Scripts and tools for optimizing quantizations in llama.cpp with GGUF imatrices.☆19Jan 10, 2025Updated last year
- Spec-driven iterative development companion CLI for OpenCode.☆17Jun 7, 2026Updated last month
- A TTS model capable of generating ultra-realistic dialogue in one pass.☆32May 1, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A bare-bones GUI application for the local inference engine, llama.cpp. Built-in TPE optimiser to find the best flags for your system☆17Jul 3, 2026Updated 2 weeks ago
- A local-first web search agent