Fast LLM swapping with sleep/wake support, compatible with vllm, llama.cpp, etc. llama-swap fork.
☆51Apr 5, 2026Updated 3 months ago
Alternatives and similar repositories for llmsnap
Users that are interested in llmsnap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An extension to oobabooga's TextGen allowing you to receive pics generated by Automatic1111's SD API☆12May 16, 2023Updated 3 years ago
- Dynamic LLM model swapping system with Docker, vLLM integration, and GPU acceleration. Supports GGUF & Hugging Face models with automatic…☆22Mar 6, 2026Updated 4 months ago
- ☆12Feb 27, 2020Updated 6 years ago
- llama-swap + a minimal ollama compatible api☆60May 26, 2026Updated 2 months ago
- Three inference engines for Llama 3: pure C for desktop systems, pure JavaScript for Node.js, and pure JavaScript for Web environments.☆25Jul 11, 2026Updated 2 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- An ansible role to install authentik, an open-source Identity Provider☆17Updated this week
- Scripts and tools for optimizing quantizations in llama.cpp with GGUF imatrices.☆19Jan 10, 2025Updated last year
- A revolutionary AI-powered framework for RimWorld that brings Large Language Models directly into colony management. Features intelligen…☆21Sep 1, 2025Updated 10 months ago
- ☆16Dec 16, 2024Updated last year
- Static analysis toolkit for LLM agent plans☆13Aug 9, 2025Updated 11 months ago
- This tool helps you abstract away common or repeated chunks of code by finding similar pieces☆15Jul 22, 2020Updated 6 years ago
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆24Jan 30, 2026Updated 6 months ago
- A proxy that hosts multiple single-model runners such as LLama.cpp and vLLM☆12May 30, 2025Updated last year
- GPU monitor for Linux terminal supporting single or multiple gpu's in realtime☆19Jul 19, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A Python-based voice assistant integrating speech-to-text (STT), text-to-speech (TTS), and powerful AI capabilities using either a local …☆19Jul 22, 2026Updated last week
- Implements harmful/harmless refusal removal using pure HF Transformers☆23May 8, 2025Updated last year
- Simple node proxy for llama-server that enables MCP use☆19May 10, 2025Updated last year
- Some of my tools for paperless-ngx, for example title generation☆11Jul 10, 2024Updated 2 years ago
- 🔥 Alternative to Ollama — multi-model serving with sub-ms model switching · CPU-only 20B inference for Edge AI · llama.cpp + stablediffu…☆39Updated this week
- Testbench for llama.cpp llama-server☆15Aug 20, 2025Updated 11 months ago
- Since the owner of the repo took it down and it used an MIT license, I guess it's okay to upload it here for people to use.☆55Mar 11, 2025Updated last year
- 🔍📃 LLM-powered PDF Table Extractor☆19Jun 26, 2025Updated last year
- Custom Players for the Slinger project☆10Oct 27, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- An open-source AI agent that lives in your terminal, privacy oriented, with no telemetry.☆74Updated this week
- Make Qwen3 Think like Gemini 2.5 Pro | Open webui function☆25May 10, 2025Updated last year
- Zoof is a high-efficiency Small Language Model (SLM) engineered from scratch. It demonstrates how modern architectural choices and high-q…☆47Jan 13, 2026Updated 6 months ago
- Medical records you can copy and paste☆12Mar 3, 2023Updated 3 years ago
- ☆15Mar 18, 2026Updated 4 months ago
- Ubiquité : Open-source Perplexity clone with multi-LLM support and KaTeX math rendering.☆48Nov 14, 2025Updated 8 months ago
- ☆18Jul 1, 2025Updated last year
- Smart OpenAI‑compatible proxy for llama.cpp: manages slots, saves/restores KV cache to disk, routes requests by prefix similarity, and pr…☆49Nov 14, 2025Updated 8 months ago
- A high-throughput and memory-efficient inference and serving engine for LLMs - Optimized for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI…☆77Jun 23, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆22Apr 3, 2026Updated 3 months ago
- Statistics for MINC volumes: A library to integrate voxel-based statistics for MINC volumes into the R environment. Supports getting and …☆26Jul 23, 2026Updated last week
- Play YouTube videos directly in your terminal with synchronized audio using ASCII rendering or ANSI truecolor.☆16Updated this week
- MusMorph, a database of standardized mouse morphology data for morphometric meta-analyses.☆12Jul 25, 2024Updated 2 years ago
- KernelSU module to fill target.txt of Tricky Store with Web UI. Also auto check if keybox valid. Your KernelSU Manager must support Web U…☆13Nov 3, 2024Updated last year
- An MCP-enabled Qwen3 0.6B demo with adjustable thinking budget, all in your browser!☆28Jun 2, 2025Updated last year
- ☆24Dec 29, 2025Updated 7 months ago