Smart OpenAI‑compatible proxy for llama.cpp: manages slots, saves/restores KV cache to disk, routes requests by prefix similarity, and protects hot slots from being overwritten. Accelerates long prompts (30–60k tokens) via instant reuse or fast on‑demand restore; supports SSE streaming and non‑stream JSON over /v1/chat/completions.
☆48Nov 14, 2025Updated 8 months ago
Alternatives and similar repositories for proxycache
Users that are interested in proxycache are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Make Qwen3 Think like Gemini 2.5 Pro | Open webui function☆25May 10, 2025Updated last year
- LLM backed Fantasy Tribe Game☆19Nov 21, 2024Updated last year
- Unified management and routing for llama.cpp, MLX and vLLM models with web dashboard.☆137Updated this week
- Generate a llama-quantize command to copy the quantization parameters of any GGUF☆34Apr 20, 2026Updated 3 months ago
- 💜 The slightly more compromising Python code formatter☆11Feb 26, 2021Updated 5 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- llama.cpp fork with additional SOTA quants and improved performance☆22Jul 10, 2026Updated 2 weeks ago
- Turbo Lossless - 1.33x Smaller, 2.93x Faster, Decode with 1 ADD operation☆15Apr 4, 2026Updated 3 months ago
- An example repo full of Douglas Adams quotes☆16Mar 12, 2024Updated 2 years ago
- The minimal AI agent engine☆39Jun 13, 2026Updated last month
- A Kubernetes controller and webhook implementation that enables safe, staged rollouts of DaemonSets☆15Jul 31, 2025Updated 11 months ago
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆18Mar 20, 2026Updated 4 months ago
- Enhancing LLMs with LoRA☆224Oct 20, 2025Updated 9 months ago
- Branch-Thinking MCP Tool A TypeScript-powered MCP server for managing parallel branches of thought, semantic cross-references, and persis…☆15Apr 25, 2025Updated last year
- Yet Another (LLM) Web UI, made with Gemini☆12Dec 25, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Deploying full-stack on-prem deep research agent that can be run entirely on a local machine for $0!☆34Nov 8, 2025Updated 8 months ago
- A high performance authentication and access-control gateway for LLM API backends☆16Jun 11, 2026Updated last month
- Home-enthusiast's guide to fine-tuning 27B+ LLMs on AMD Strix Halo (gfx1151, Ryzen AI MAX+ 395) — the patches and tuning to make Linux ma…☆26Jun 9, 2026Updated last month
- VellumForge2 is a Golang CLI for generating high-quality Direct Preference Optimization datasets via a hierarchical prompt pipeline with …☆20Feb 15, 2026Updated 5 months ago
- ☆22May 12, 2026Updated 2 months ago
- How much of your code is written by agents?☆23Apr 24, 2026Updated 3 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆2,967Updated this week
- ☆15Feb 10, 2026Updated 5 months ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,152Updated this week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Codex-style multi-agent orchestration for pi☆21Mar 26, 2026Updated 4 months ago
- Modern Memory Bandwidth and Latency Benchmarks☆18Jun 10, 2026Updated last month
- 🎙️ VibeVoice FastAPI - Multi-Speaker TTS API☆32Aug 27, 2025Updated 10 months ago
- AI in A Box☆28Updated this week
- FamilyBench evaluation tool for testing the relational reasoning capabilities of Large Language Models (LLMs).☆47May 4, 2026Updated 2 months ago
- ☆16Oct 20, 2024Updated last year
- Deploy Sourcegraph on Kubernetes using Helm☆18Updated this week
- Push Notification Relay Server for Frappe Apps☆12Dec 24, 2025Updated 7 months ago
- HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE sp…☆22Jun 23, 2026Updated last month
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆11Feb 20, 2025Updated last year
- Working with LLM in C#☆16Jul 14, 2026Updated last week
- Desktop application for instant AI-powered text transformation. Translate, correct, summarize, and change the tone of any text, anywhere,…☆35Dec 29, 2025Updated 6 months ago
- Serving LLMs in the HF-Transformers format via a PyFlask API☆72Sep 10, 2024Updated last year
- A physics-grounded, cost-aware optimizer for vLLM.☆56Updated this week
- Evolutionary strategies finetuning library for LLMs☆25Jun 29, 2026Updated 3 weeks ago
- Linux & Powershell scripts to easily set up and run the Qwen 3.5 series locally on Windows and Linux with llama.cpp.☆92Apr 28, 2026Updated 2 months ago