Smart OpenAI‑compatible proxy for llama.cpp: manages slots, saves/restores KV cache to disk, routes requests by prefix similarity, and protects hot slots from being overwritten. Accelerates long prompts (30–60k tokens) via instant reuse or fast on‑demand restore; supports SSE streaming and non‑stream JSON over /v1/chat/completions.
☆55Sep 16, 2026Updated last week
Alternatives and similar repositories for proxycache
Users that are interested in proxycache are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Make Qwen3 Think like Gemini 2.5 Pro | Open webui function☆25May 10, 2025Updated last year
- BROKEN REPO. DO NOT USE UNDER ANY CIRCUMSTANCES☆20Sep 21, 2026Updated last week
- Unified management and routing for llama.cpp, MLX and vLLM models with web dashboard.☆154Updated this week
- Generate a llama-quantize command to copy the quantization parameters of any GGUF☆36Apr 20, 2026Updated 5 months ago
- 💜 The slightly more compromising Python code formatter☆11Feb 26, 2021Updated 5 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Agentic BYOK Browser-Based Website Builder☆61Updated this week
- llama.cpp fork with additional SOTA quants and improved performance☆22Updated this week
- ☆59Oct 10, 2025Updated 11 months ago
- Turbo Lossless - 1.33x Smaller, 2.93x Faster, Decode with 1 ADD operation☆15Apr 4, 2026Updated 5 months ago
- Proxy for OpenAI☆16Sep 2, 2025Updated last year
- Private-first, self-hostable knowledge base. Your data, your server, your control. No cloud, no telemetry, no trust required.☆22Aug 15, 2026Updated last month
- btop-style terminal dashboards for Prometheus & Grafana.☆51Sep 7, 2026Updated 2 weeks ago
- A Kubernetes controller and webhook implementation that enables safe, staged rollouts of DaemonSets☆16Jul 31, 2025Updated last year
- stt for home assistant [gigaam, vosk[ru], parakeet, canary]☆26Jul 16, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Branch-Thinking MCP Tool A TypeScript-powered MCP server for managing parallel branches of thought, semantic cross-references, and persis…☆15Apr 25, 2025Updated last year
- Enhancing LLMs with LoRA☆227Oct 20, 2025Updated 11 months ago
- A teminal othello (reversi) in Nim.☆10May 2, 2021Updated 5 years ago
- ☆16Feb 5, 2026Updated 7 months ago
- Yet Another (LLM) Web UI, made with Gemini☆12Dec 25, 2024Updated last year
- A high performance authentication and access-control gateway for LLM API backends☆16Jun 11, 2026Updated 3 months ago
- Home-enthusiast's guide to fine-tuning 27B+ LLMs on AMD Strix Halo (gfx1151, Ryzen AI MAX+ 395) — the patches and tuning to make Linux ma…☆29Jun 9, 2026Updated 3 months ago
- For a number of years now, work has been proceeding in order to bring to perfection the crudely-conceived idea of a machine that would no…☆14Nov 12, 2025Updated 10 months ago
- VellumForge2 is a Golang CLI for generating high-quality Direct Preference Optimization datasets via a hierarchical prompt pipeline with …☆22Feb 15, 2026Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆23Apr 3, 2026Updated 5 months ago
- ☆27May 12, 2026Updated 4 months ago
- ☆19Updated this week
- Deploying full-stack on-prem deep research agent that can be run entirely on a local machine for $0!☆34Nov 8, 2025Updated 10 months ago
- How much of your code is written by agents?☆24Apr 24, 2026Updated 5 months ago
- Reference demos for Predicate SDK AgentRuntime and Debugger for per-step & task verifications☆23Apr 11, 2026Updated 5 months ago
- ☆45Oct 9, 2025Updated 11 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,260Updated this week
- ☆33Apr 12, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆15Feb 10, 2026Updated 7 months ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,758Updated this week
- Codex-style multi-agent orchestration for pi☆21Mar 26, 2026Updated 6 months ago
- 🎙️ VibeVoice FastAPI - Multi-Speaker TTS API☆33Aug 27, 2025Updated last year
- Modern Memory Bandwidth and Latency Benchmarks☆18Aug 27, 2026Updated last month
- unix domain sockets that look just like tcp sockets☆11Jun 21, 2018Updated 8 years ago
- AI in A Box☆28Aug 21, 2026Updated last month