Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale ππ¦ Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.
β1,643Jul 19, 2026Updated this week
Alternatives and similar repositories for paddler
Users that are interested in paddler are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Fast, flexible LLM inferenceβ7,508Updated this week
- β615Updated this week
- Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etcβ5,087Updated this week
- Extracts structured data from unstructured input. Programming language agnostic. Uses llama.cppβ45May 16, 2024Updated 2 years ago
- llama.cpp fork with additional SOTA quants and improved performanceβ2,943Updated this week
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.β3,004Jul 5, 2026Updated 2 weeks ago
- Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrailsβ¦β6,882Updated this week
- Distribute and run LLMs with a single file.β25,416Updated this week
- WebAssembly binding for llama.cpp - Enabling on-browser LLM inferenceβ1,151Jun 17, 2026Updated last month
- Minimalist ML framework for Rustβ20,696Jul 14, 2026Updated last week
- LLM inference in C/C++β121,178Updated this week
- Pure C++ implementation of several models for real-time chatting on your computer (CPU & GPU)β912Updated this week
- Tensor library for machine learningβ15,034Updated this week
- Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++β6,560Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inβ¦β2,969Updated this week
- Large-scale LLM inference engineβ1,806Updated this week
- Optimizing inference proxy for LLMsβ4,182Updated this week
- Inference at the speed of light.β2,888Updated this week
- Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.β700Jul 14, 2026Updated last week
- [Unmaintained, see README] An ecosystem of Rust libraries for working with large language modelsβ6,153Jun 24, 2024Updated 2 years ago
- Qdrant - High-performance, massive-scale Vector Database and Vector Search Engine for the next generation of AI. Also available in the clβ¦β33,466Updated this week
- A fast inference library for running LLMs locally on modern consumer-class GPUsβ4,586Mar 4, 2026Updated 4 months ago
- Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.β15,612Updated this week
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A vector search SQLite extension that runs anywhere!β7,915May 18, 2026Updated 2 months ago
- π‘ All-in-one AI framework for semantic search, LLM orchestration and language model workflowsβ12,741Updated this week
- Structured Outputsβ14,833Updated this week
- A high-performance inference engine for AI modelsβ1,659Updated this week
- Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuildβ3,905Updated this week
- Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data veβ¦β6,832Updated this week
- Fast ML inference & training for ONNX models in Rustβ2,412Updated this week
- A SQL database in Rust: SQLite-compatible, now also speaking Postgres (experimental). The LLVM of databases.β23,283Updated this week
- π§± easy, fast and local-first microVM runtimeβ6,987Updated this week
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- βοΈπ¦ Build modular and scalable LLM Applications in Rustβ7,992Updated this week
- The llama-cpp-agent framework is a tool designed for easy interaction with Large Language Models (LLMs). Allowing users to chat with LLM β¦β650Mar 9, 2026Updated 4 months ago
- Unsloth is a local UI for training and running Gemma 4, Qwen3.6, DeepSeek, Kimi, GLM and other models.β68,666Updated this week
- Go ahead and axolotl questionsβ12,222Updated this week
- Self-hosted AI coding assistantβ33,779Jun 30, 2026Updated 3 weeks ago
- Developer-friendly OSS embedded retrieval library for multimodal AI. Search More; Manage Less.β10,945Updated this week
- VS Code extension for LLM-assisted code/text completionβ1,451Updated this week