Smart OpenAI‑compatible proxy for llama.cpp: manages slots, saves/restores KV cache to disk, routes requests by prefix similarity, and protects hot slots from being overwritten. Accelerates long prompts (30–60k tokens) via instant reuse or fast on‑demand restore; supports SSE streaming and non‑stream JSON over /v1/chat/completions.
☆51Nov 14, 2025Updated 9 months ago
Alternatives and similar repositories for proxycache
Users that are interested in proxycache are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Make Qwen3 Think like Gemini 2.5 Pro | Open webui function☆25May 10, 2025Updated last year
- Unified management and routing for llama.cpp, MLX and vLLM models with web dashboard.☆142Aug 11, 2026Updated last week
- A PyTorch framework for training transformer language models with Mixture of Experts (MoE) architecture support, Mixture of Depths (MoD),…☆21Aug 10, 2026Updated last week
- Generate a llama-quantize command to copy the quantization parameters of any GGUF☆36Apr 20, 2026Updated 3 months ago
- Agentic BYOK Browser-Based Website Builder☆54Updated this week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- llama.cpp fork with additional SOTA quants and improved performance☆22Updated this week
- ☆58Oct 10, 2025Updated 10 months ago
- CLI secret management☆16Updated this week
- Proxy for OpenAI☆16Sep 2, 2025Updated 11 months ago
- An example repo full of Douglas Adams quotes☆16Mar 12, 2024Updated 2 years ago
- Private-first, self-hostable knowledge base. Your data, your server, your control. No cloud, no telemetry, no trust required.☆22Updated this week
- The minimal AI agent engine☆39Aug 1, 2026Updated 2 weeks ago
- ☆17Jun 4, 2026Updated 2 months ago
- grom — btop-style terminal dashboards for Prometheus & Grafana☆16Aug 6, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A Kubernetes controller and webhook implementation that enables safe, staged rollouts of DaemonSets☆16Jul 31, 2025Updated last year
- Enhancing LLMs with LoRA☆225Oct 20, 2025Updated 9 months ago
- Branch-Thinking MCP Tool A TypeScript-powered MCP server for managing parallel branches of thought, semantic cross-references, and persis…☆15Apr 25, 2025Updated last year
- Yet Another (LLM) Web UI, made with Gemini☆12Dec 25, 2024Updated last year
- ☆18Jan 5, 2016Updated 10 years ago
- Deploying full-stack on-prem deep research agent that can be run entirely on a local machine for $0!☆34Nov 8, 2025Updated 9 months ago
- A high performance authentication and access-control gateway for LLM API backends☆16Jun 11, 2026Updated 2 months ago
- CLI tool for bundling project files as LLM context.☆50Jul 29, 2026Updated 2 weeks ago
- For a number of years now, work has been proceeding in order to bring to perfection the crudely-conceived idea of a machine that would no…☆14Nov 12, 2025Updated 9 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆22Apr 3, 2026Updated 4 months ago
- ☆18Updated this week
- How much of your code is written by agents?☆24Apr 24, 2026Updated 3 months ago
- Reference demos for Predicate SDK AgentRuntime and Debugger for per-step & task verifications☆23Apr 11, 2026Updated 4 months ago
- ☆43Oct 9, 2025Updated 10 months ago
- ☆15Feb 10, 2026Updated 6 months ago
- ☆33Apr 12, 2026Updated 4 months ago
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc☆5,393Updated this week
- Modern Memory Bandwidth and Latency Benchmarks☆18Jun 10, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- FamilyBench evaluation tool for testing the relational reasoning capabilities of Large Language Models (LLMs).☆46May 4, 2026Updated 3 months ago
- AI in A Box☆28Jul 26, 2026Updated 3 weeks ago
- A Grafana data source plugin that brings semantic layer analytics via Cube. Define metrics once, use them consistently across all dashboa…☆18Updated this week
- Deploy Sourcegraph on Kubernetes using Helm☆18Aug 11, 2026Updated last week
- HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE sp…☆23Jun 23, 2026Updated last month
- Serving LLMs in the HF-Transformers format via a PyFlask API☆72Sep 10, 2024Updated last year
- Desktop application for instant AI-powered text transformation. Translate, correct, summarize, and change the tone of any text, anywhere,…☆35Dec 29, 2025Updated 7 months ago