Fast LLM swapping with sleep/wake support, compatible with vllm, llama.cpp, etc. llama-swap fork.
☆56Apr 5, 2026Updated 5 months ago
Alternatives and similar repositories for llmsnap
Users that are interested in llmsnap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- vLLM fork with Marlin W4A8 SM121 patches + TMA module☆15Mar 23, 2026Updated 5 months ago
- An extension to oobabooga's TextGen allowing you to receive pics generated by Automatic1111's SD API☆11May 16, 2023Updated 3 years ago
- ☆12Feb 27, 2020Updated 6 years ago
- Generate Your Own Private Morning Radio for Commute☆33Feb 5, 2025Updated last year
- Web application for fine-tuning language models☆15Aug 16, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Ansible role for installing authentik☆17Updated this week
- Scripts and tools for optimizing quantizations in llama.cpp with GGUF imatrices.☆19Jan 10, 2025Updated last year
- A revolutionary AI-powered framework for RimWorld that brings Large Language Models directly into colony management. Features intelligen…☆23Sep 1, 2025Updated last year
- ☆16Dec 16, 2024Updated last year
- TeamCopilot: Deploy AI agents for your team to automate business workflows and coding.☆15Jun 23, 2026Updated 2 months ago
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆25Jan 30, 2026Updated 7 months ago
- A proxy that hosts multiple single-model runners such as LLama.cpp and vLLM☆12May 30, 2025Updated last year
- GPU monitor for Linux terminal supporting single or multiple gpu's in realtime☆20Jul 19, 2026Updated last month
- A simple, observable code-writing agent builder in TypeScript.☆33Apr 9, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- A Python-based voice assistant integrating speech-to-text (STT), text-to-speech (TTS), and powerful AI capabilities using either a local …☆19Jul 22, 2026Updated last month
- Implements harmful/harmless refusal removal using pure HF Transformers☆25May 8, 2025Updated last year
- Simple node proxy for llama-server that enables MCP use☆19May 10, 2025Updated last year
- 🔥 Alternative to Ollama — multi-model serving with sub-ms model switching · CPU-only 20B inference for Edge AI · llama.cpp + stablediffu…☆40Aug 22, 2026Updated 3 weeks ago
- Some of my tools for paperless-ngx, for example title generation☆11Jul 10, 2024Updated 2 years ago
- Install guide for putting Debian GNU/Linux on a PogoPlug Pro☆10Jan 19, 2023Updated 3 years ago
- Testbench for llama.cpp llama-server☆15Aug 20, 2025Updated last year
- Since the owner of the repo took it down and it used an MIT license, I guess it's okay to upload it here for people to use.☆55Mar 11, 2025Updated last year
- Superseded by github.com/poisonxa16/pxa — PXA, the set-and-forget engine for Pascal and Volta☆24Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 🔍📃 LLM-powered PDF Table Extractor☆19Jun 26, 2025Updated last year
- Custom Players for the Slinger project☆10Oct 27, 2024Updated last year
- An open-source AI agent that lives in your terminal, privacy oriented, with no telemetry.☆79Updated this week
- Make Qwen3 Think like Gemini 2.5 Pro | Open webui function☆25May 10, 2025Updated last year
- Zoof is a high-efficiency Small Language Model (SLM) engineered from scratch. It demonstrates how modern architectural choices and high-q…☆47Jan 13, 2026Updated 7 months ago
- ☆15Mar 18, 2026Updated 5 months ago
- Smooth is an SDK and a CLI for AI browser automation☆19Aug 14, 2026Updated 3 weeks ago
- Ubiquité : Open-source Perplexity clone with multi-LLM support and KaTeX math rendering.☆48Nov 14, 2025Updated 9 months ago
- DeepSeek-V4-Flash on Ampere SM 8.6 via vLLM (pyref kernel replacements)☆35Jul 2, 2026Updated 2 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Constellation Network Node Administration Utility☆20Jan 18, 2024Updated 2 years ago
- Smart OpenAI‑compatible proxy for llama.cpp: manages slots, saves/restores KV cache to disk, routes requests by prefix similarity, and pr…☆53Nov 14, 2025Updated 9 months ago
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆23Apr 3, 2026Updated 5 months ago
- AI agent rules: markdown files for Claude.md, ChatGPT, Copilot, Cursor, Windsurf, and more.☆25Aug 30, 2026Updated last week
- An MCP-enabled Qwen3 0.6B demo with adjustable thinking budget, all in your browser!☆28Jun 2, 2025Updated last year
- Template for a new generative ai project using uv, nicegui, fastapi, llms (cloud & local with litellm and ollama, cpu/gpu) and langfuse f…☆118Updated this week
- Hunterx | GET YUOR MAX PERFORMANCE | Module Magisk☆15Aug 11, 2023Updated 3 years ago