Fast LLM swapping with sleep/wake support, compatible with vllm, llama.cpp, etc. llama-swap fork.
☆57Apr 5, 2026Updated 5 months ago
Alternatives and similar repositories for llmsnap
Users that are interested in llmsnap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An extension to oobabooga's TextGen allowing you to receive pics generated by Automatic1111's SD API☆11May 16, 2023Updated 3 years ago
- Dynamic LLM model swapping system with Docker, vLLM integration, and GPU acceleration. Supports GGUF & Hugging Face models with automatic…☆23Mar 6, 2026Updated 6 months ago
- Generate Your Own Private Morning Radio for Commute☆33Feb 5, 2025Updated last year
- llama-swap + a minimal ollama compatible api☆63Updated this week
- Ansible role for installing authentik☆17Sep 24, 2026Updated last week
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Scripts and tools for optimizing quantizations in llama.cpp with GGUF imatrices.☆19Jan 10, 2025Updated last year
- A revolutionary AI-powered framework for RimWorld that brings Large Language Models directly into colony management. Features intelligen…☆24Sep 1, 2025Updated last year
- ☆16Dec 16, 2024Updated last year
- Static analysis toolkit for LLM agent plans☆13Aug 9, 2025Updated last year
- This tool helps you abstract away common or repeated chunks of code by finding similar pieces☆15Jul 22, 2020Updated 6 years ago
- Multi-GPU device selection for LTXV2 video generation in ComfyUI☆34Jan 10, 2026Updated 8 months ago
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆27Jan 30, 2026Updated 8 months ago
- A proxy that hosts multiple single-model runners such as LLama.cpp and vLLM☆12May 30, 2025Updated last year
- Documenting all the concepts and codes learned in NLP☆29Sep 19, 2021Updated 5 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- GPU monitor for Linux terminal supporting single or multiple gpu's in realtime☆20Jul 19, 2026Updated 2 months ago
- A simple, observable code-writing agent builder in TypeScript.☆33Apr 9, 2025Updated last year
- A Python-based voice assistant integrating speech-to-text (STT), text-to-speech (TTS), and powerful AI capabilities using either a local …☆19Jul 22, 2026Updated 2 months ago
- ☆65Updated this week
- Implements harmful/harmless refusal removal using pure HF Transformers☆25May 8, 2025Updated last year
- 🔥 Alternative to Ollama — multi-model serving with sub-ms model switching · CPU-only 20B inference for Edge AI · llama.cpp + stablediffu…☆40Updated this week
- Simple node proxy for llama-server that enables MCP use☆19May 10, 2025Updated last year
- Some of my tools for paperless-ngx, for example title generation☆11Jul 10, 2024Updated 2 years ago
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆133Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- PromptMII: Meta-Learning Instruction Induction for LLMs☆48Jan 12, 2026Updated 8 months ago
- Testbench for llama.cpp llama-server☆15Aug 20, 2025Updated last year
- Since the owner of the repo took it down and it used an MIT license, I guess it's okay to upload it here for people to use.☆55Mar 11, 2025Updated last year
- Superseded by github.com/poisonxa16/pxa — PXA, the set-and-forget engine for Pascal and Volta☆24Sep 8, 2026Updated 3 weeks ago
- 🔍📃 LLM-powered PDF Table Extractor☆19Jun 26, 2025Updated last year
- Custom Players for the Slinger project☆10Oct 27, 2024Updated last year
- An open-source AI agent that lives in your terminal, privacy oriented, with no telemetry.☆82Updated this week
- Make Qwen3 Think like Gemini 2.5 Pro | Open webui function☆25May 10, 2025Updated last year
- Zoof is a high-efficiency Small Language Model (SLM) engineered from scratch. It demonstrates how modern architectural choices and high-q…☆48Jan 13, 2026Updated 8 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Install micromamba, and optionally create a base conda environment.☆10Apr 5, 2025Updated last year
- ☆15Mar 18, 2026Updated 6 months ago
- Ubiquité : Open-source Perplexity clone with multi-LLM support and KaTeX math rendering.☆49Nov 14, 2025Updated 10 months ago
- ☆18Jul 1, 2025Updated last year
- A high-throughput and memory-efficient inference and serving engine for LLMs - Optimized for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI…☆91Jun 23, 2026Updated 3 months ago
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆24Apr 3, 2026Updated 5 months ago
- vLLM fork for Tesla V100 (SM70) — extends 1CatAI's AWQ support and adds GGUF support☆23Jun 20, 2026Updated 3 months ago