Fast LLM swapping with sleep/wake support, compatible with vllm, llama.cpp, etc. llama-swap fork.
☆53Apr 5, 2026Updated 4 months ago
Alternatives and similar repositories for llmsnap
Users that are interested in llmsnap are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Python SDK for Agent Vector Protocol – transfer KV-cache between LLM agents instead of text☆26Apr 26, 2026Updated 3 months ago
- llama-swap + a minimal ollama compatible api☆61May 26, 2026Updated 2 months ago
- An ansible role to install authentik, an open-source Identity Provider☆17Updated this week
- Scripts and tools for optimizing quantizations in llama.cpp with GGUF imatrices.☆19Jan 10, 2025Updated last year
- A revolutionary AI-powered framework for RimWorld that brings Large Language Models directly into colony management. Features intelligen…☆22Sep 1, 2025Updated 11 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆30Updated this week
- TeamCopilot: Deploy AI agents for your team to automate business workflows and coding.☆15Jun 23, 2026Updated last month
- IMMA-based **FP8-as-storage** GEMM experiments for Ampere (sm_86 / RTX 3090 Ti).☆24Jan 30, 2026Updated 6 months ago
- A proxy that hosts multiple single-model runners such as LLama.cpp and vLLM☆12May 30, 2025Updated last year
- A practical, hands-on guide to building a small language model from scratch. Learn transformer architecture, attention mechanisms, and tr…☆20Dec 10, 2025Updated 8 months ago
- Repo for YouTube tutorial on how to self-host a MinIo instance and connect a Next.js 15 app to MinIo and upload files☆20Jan 15, 2025Updated last year
- GPU monitor for Linux terminal supporting single or multiple gpu's in realtime☆19Jul 19, 2026Updated 3 weeks ago
- A simple, observable code-writing agent builder in TypeScript.☆33Apr 9, 2025Updated last year
- A Python-based voice assistant integrating speech-to-text (STT), text-to-speech (TTS), and powerful AI capabilities using either a local …☆19Jul 22, 2026Updated 3 weeks ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- PXQ: PXA-native low-bit MoE quants (2/3/4-bit, E16-row scales) + fused CUDA kernels for Pascal/Volta — run a real 35B on a salvaged 12-16…☆15Aug 11, 2026Updated last week
- Implements harmful/harmless refusal removal using pure HF Transformers☆25May 8, 2025Updated last year
- Simple node proxy for llama-server that enables MCP use☆19May 10, 2025Updated last year
- 🔥 Alternative to Ollama — multi-model serving with sub-ms model switching · CPU-only 20B inference for Edge AI · llama.cpp + stablediffu…☆39Updated this week
- Some of my tools for paperless-ngx, for example title generation☆11Jul 10, 2024Updated 2 years ago
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆132Aug 12, 2026Updated last week
- PromptMII: Meta-Learning Instruction Induction for LLMs☆48Jan 12, 2026Updated 7 months ago
- Testbench for llama.cpp llama-server☆15Aug 20, 2025Updated 11 months ago
- Since the owner of the repo took it down and it used an MIT license, I guess it's okay to upload it here for people to use.☆55Mar 11, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 🔍📃 LLM-powered PDF Table Extractor☆19Jun 26, 2025Updated last year
- An open-source AI agent that lives in your terminal, privacy oriented, with no telemetry.☆78Updated this week
- Make Qwen3 Think like Gemini 2.5 Pro | Open webui function☆25May 10, 2025Updated last year
- Hunterx | GET YUOR MAX PERFORMANCE | Module Magisk☆14Aug 11, 2023Updated 3 years ago
- Install micromamba, and optionally create a base conda environment.☆10Apr 5, 2025Updated last year
- ☆15Mar 18, 2026Updated 5 months ago
- Smooth is an SDK and a CLI for AI browser automation☆19Updated this week
- ☆18Jul 1, 2025Updated last year
- vLLM fork for Tesla V100 (SM70) — extends 1CatAI's AWQ support and adds GGUF support☆21Jun 20, 2026Updated last month
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Play YouTube videos directly in your terminal with synchronized audio using ASCII rendering or ANSI truecolor.☆17Jul 23, 2026Updated 3 weeks ago
- AI agent rules: markdown files for Claude.md, ChatGPT, Copilot, Cursor, Windsurf, and more.☆25Jul 12, 2026Updated last month
- An MCP-enabled Qwen3 0.6B demo with adjustable thinking budget, all in your browser!☆28Jun 2, 2025Updated last year
- Template for a new generative ai project using uv, nicegui, fastapi, llms (cloud & local with litellm and ollama, cpu/gpu) and langfuse f…☆119Updated this week
- Analyze Reddit posts☆32Jun 5, 2026Updated 2 months ago
- Stop prompting your agents to behave. Start engineering them to.☆25Apr 17, 2026Updated 4 months ago
- A simple push notification gateway for Rocket.Chat servers written in Go.☆14Nov 20, 2025Updated 8 months ago