Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale ππ¦ Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.
β1,676Sep 17, 2026Updated this week
Alternatives and similar repositories for paddler
Users that are interested in paddler are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Fast, flexible LLM inferenceβ7,716Sep 8, 2026Updated 2 weeks ago
- β650Updated this week
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etcβ5,722Updated this week
- llama.cpp fork with additional SOTA quants and improved performanceβ3,254Updated this week
- Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.β3,061Jul 5, 2026Updated 2 months ago
- End-to-end encrypted cloud storage - Proton Drive β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrailsβ¦β7,062Aug 19, 2026Updated last month
- Distribute and run LLMs with a single file.β26,028Updated this week
- WebAssembly binding for llama.cpp - Enabling on-browser LLM inferenceβ1,288Updated this week
- Minimalist ML framework for Rustβ21,087Updated this week
- LLM inference in C/C++β129,197Updated this week
- Pure C++ implementation of several models for real-time chatting on your computer (CPU & GPU)β930Sep 7, 2026Updated 2 weeks ago
- Tensor library for machine learningβ15,390Updated this week
- Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++β7,092Updated this week
- Structured Outputsβ15,873Updated this week
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inβ¦β3,056Updated this week
- Large-scale LLM inference engineβ1,865Sep 11, 2026Updated last week
- Optimizing inference proxy for LLMsβ4,303Updated this week
- Inference at the speed of light.β2,990Updated this week
- A vector search SQLite extension that runs anywhere!β8,127May 18, 2026Updated 4 months ago
- Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.β728Updated this week
- [Unmaintained, see README] An ecosystem of Rust libraries for working with large language modelsβ6,156Jun 24, 2024Updated 2 years ago
- Qdrant - High-performance, massive-scale Vector Database and Vector Search Engine for the next generation of AI. Also available in the clβ¦β34,754Updated this week
- A fast inference library for running LLMs locally on modern consumer-class GPUsβ4,626Mar 4, 2026Updated 6 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.β15,965Updated this week
- A high-performance inference engine for AI modelsβ1,810Updated this week
- π‘ All-in-one AI framework for semantic search, LLM orchestration and language model workflowsβ12,966Updated this week
- Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuildβ4,075Updated this week
- Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data veβ¦β7,106Updated this week
- Fast ML inference & training for ONNX models in Rustβ2,523Updated this week
- π§± fast branchable micro virtual machines for any workloadβ8,385Updated this week
- A SQL database in Rust: SQLite-compatible, now also speaking Postgres (experimental). The LLVM of databases.β24,347Updated this week
- The llama-cpp-agent framework is a tool designed for easy interaction with Large Language Models (LLMs). Allowing users to chat with LLM β¦β658Mar 9, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- βοΈπ¦ Build modular and scalable LLM Applications in Rustβ8,698Updated this week
- Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.β76,575Updated this week
- Go ahead and axolotl questionsβ12,491Updated this week
- Self-hosted AI coding assistantβ33,889Jun 30, 2026Updated 2 months ago
- Developer-friendly OSS embedded retrieval library for multimodal AI. Search More; Manage Less.β11,503Updated this week
- Apache Iggy: Hyper-Efficient Message Streaming at Laser Speedβ4,926Updated this week
- VS Code extension for LLM-assisted code/text completionβ1,516Updated this week