Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale ππ¦ Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.
β1,665Jul 19, 2026Updated last month
Alternatives and similar repositories for paddler
Users that are interested in paddler are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Fast, flexible LLM inferenceβ7,639Updated this week
- β639Updated this week
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etcβ5,527Updated this week
- Extracts structured data from unstructured input. Programming language agnostic. Uses llama.cppβ45May 16, 2024Updated 2 years ago
- llama.cpp fork with additional SOTA quants and improved performanceβ3,160Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.β3,049Jul 5, 2026Updated last month
- Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrailsβ¦β7,028Aug 19, 2026Updated last week
- Distribute and run LLMs with a single file.β25,819Updated this week
- WebAssembly binding for llama.cpp - Enabling on-browser LLM inferenceβ1,189Updated this week
- Minimalist ML framework for Rustβ20,977Updated this week
- LLM inference in C/C++β126,516Updated this week
- Pure C++ implementation of several models for real-time chatting on your computer (CPU & GPU)β921Updated this week
- Tensor library for machine learningβ15,267Updated this week
- Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++β6,887Updated this week
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Structured Outputsβ15,726Updated this week
- RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inβ¦β3,026Updated this week
- Large-scale LLM inference engineβ1,844Aug 13, 2026Updated 2 weeks ago
- Optimizing inference proxy for LLMsβ4,259Jul 18, 2026Updated last month
- Inference at the speed of light.β2,963Updated this week
- A vector search SQLite extension that runs anywhere!β8,061May 18, 2026Updated 3 months ago
- Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.β718Aug 20, 2026Updated last week
- [Unmaintained, see README] An ecosystem of Rust libraries for working with large language modelsβ6,155Jun 24, 2024Updated 2 years ago
- Qdrant - High-performance, massive-scale Vector Database and Vector Search Engine for the next generation of AI. Also available in the clβ¦β34,294Updated this week
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A fast inference library for running LLMs locally on modern consumer-class GPUsβ4,613Mar 4, 2026Updated 5 months ago
- Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.β15,842Updated this week
- π‘ All-in-one AI framework for semantic search, LLM orchestration and language model workflowsβ12,913Updated this week
- A high-performance inference engine for AI modelsβ1,684Updated this week
- Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuildβ4,011Updated this week
- Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data veβ¦β7,010Updated this week
- Fast ML inference & training for ONNX models in Rustβ2,479Updated this week
- A SQL database in Rust: SQLite-compatible, now also speaking Postgres (experimental). The LLVM of databases.β24,102Updated this week
- π§± easy fast local-first microVM runtime and libraryβ8,034Updated this week
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The llama-cpp-agent framework is a tool designed for easy interaction with Large Language Models (LLMs). Allowing users to chat with LLM β¦β660Mar 9, 2026Updated 5 months ago
- βοΈπ¦ Build modular and scalable LLM Applications in Rustβ8,469Updated this week
- Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.β75,340Updated this week
- Go ahead and axolotl questionsβ12,430Updated this week
- Self-hosted AI coding assistantβ33,846Jun 30, 2026Updated 2 months ago
- Developer-friendly OSS embedded retrieval library for multimodal AI. Search More; Manage Less.β11,321Updated this week
- VS Code extension for LLM-assisted code/text completionβ1,495Aug 19, 2026Updated last week