Smart OpenAI‑compatible proxy for llama.cpp: manages slots, saves/restores KV cache to disk, routes requests by prefix similarity, and protects hot slots from being overwritten. Accelerates long prompts (30–60k tokens) via instant reuse or fast on‑demand restore; supports SSE streaming and non‑stream JSON over /v1/chat/completions.
☆53Nov 14, 2025Updated 9 months ago
Alternatives and similar repositories for proxycache
Users that are interested in proxycache are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Make Qwen3 Think like Gemini 2.5 Pro | Open webui function☆25May 10, 2025Updated last year
- BROKEN REPO. DO NOT USE UNDER ANY CIRCUMSTANCES☆21Aug 31, 2026Updated last week
- Unified management and routing for llama.cpp, MLX and vLLM models with web dashboard.☆150Aug 31, 2026Updated last week
- Generate a llama-quantize command to copy the quantization parameters of any GGUF☆36Apr 20, 2026Updated 4 months ago
- Agentic BYOK Browser-Based Website Builder☆59Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- llama.cpp fork with additional SOTA quants and improved performance☆22Aug 31, 2026Updated last week
- ☆59Oct 10, 2025Updated 10 months ago
- CLI secret management☆16Aug 13, 2026Updated 3 weeks ago
- An example repo full of Douglas Adams quotes☆16Mar 12, 2024Updated 2 years ago
- Private-first, self-hostable knowledge base. Your data, your server, your control. No cloud, no telemetry, no trust required.☆22Aug 15, 2026Updated 3 weeks ago
- btop-style terminal dashboards for Prometheus & Grafana.☆51Updated this week
- A Kubernetes controller and webhook implementation that enables safe, staged rollouts of DaemonSets☆16Jul 31, 2025Updated last year
- stt for home assistant [gigaam, vosk[ru], parakeet, canary]☆23Jul 16, 2026Updated last month
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆18Mar 20, 2026Updated 5 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Enhancing LLMs with LoRA☆224Oct 20, 2025Updated 10 months ago
- Yet Another (LLM) Web UI, made with Gemini☆12Dec 25, 2024Updated last year
- A high performance authentication and access-control gateway for LLM API backends☆16Jun 11, 2026Updated 2 months ago
- CLI tool for bundling project files as LLM context.☆53Aug 30, 2026Updated last week
- For a number of years now, work has been proceeding in order to bring to perfection the crudely-conceived idea of a machine that would no…☆14Nov 12, 2025Updated 9 months ago
- VellumForge2 is a Golang CLI for generating high-quality Direct Preference Optimization datasets via a hierarchical prompt pipeline with …☆21Feb 15, 2026Updated 6 months ago
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆23Apr 3, 2026Updated 5 months ago
- ☆24May 12, 2026Updated 3 months ago
- Persistent REPL scratchpad for coding agents — variables survive across turns, only print() enters context. Agent skill for Claude Code, …☆22Mar 22, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- How much of your code is written by agents?☆24Apr 24, 2026Updated 4 months ago
- ☆44Oct 9, 2025Updated 10 months ago
- llama.cpp fork with additional SOTA quants and improved performance☆3,190Updated this week
- ☆33Apr 12, 2026Updated 4 months ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,598Updated this week
- Codex-style multi-agent orchestration for pi☆21Mar 26, 2026Updated 5 months ago
- 🎙️ VibeVoice FastAPI - Multi-Speaker TTS API☆33Aug 27, 2025Updated last year
- Modern Memory Bandwidth and Latency Benchmarks☆18Aug 27, 2026Updated last week
- FamilyBench evaluation tool for testing the relational reasoning capabilities of Large Language Models (LLMs).☆46Updated this week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- AI in A Box☆28Aug 21, 2026Updated 2 weeks ago
- A Grafana data source plugin that brings semantic layer analytics via Cube. Define metrics once, use them consistently across all dashboa…☆18Updated this week
- HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE sp…☆25Jun 23, 2026Updated 2 months ago
- ☆11Feb 20, 2025Updated last year
- Mixin classes and traits dynamically☆10Sep 4, 2017Updated 9 years ago
- Serving LLMs in the HF-Transformers format via a PyFlask API☆72Sep 10, 2024Updated last year
- Desktop application for instant AI-powered text transformation. Translate, correct, summarize, and change the tone of any text, anywhere,…☆35Dec 29, 2025Updated 8 months ago