Autonomous self-improving 4x DGX Spark (GB10) MoA stack + LoRA loop (DSV4F router, Qwen3.6/Omni/TwoTower/Gemma). Hermes MoA routing, ~90% local, self-LoRA. Includes a SINGLE-SPARK recipe (all-Gemma-4: 31B brain + DiffusionGemma + 12B) to build the same on 1 GB10.
☆24Aug 17, 2026Updated last month
Alternatives and similar repositories for Keys-Setup-Autonomous-Self-Improving-Local-Inference-Stack
Users that are interested in Keys-Setup-Autonomous-Self-Improving-Local-Inference-Stack are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MiniMax-M3 (428B, no pruning) at 36 tok/s on 2× NVIDIA DGX Spark — W4A16 GPTQ + NVFP4 KV + EAGLE-3 speculative decoding on vLLM. Three se…☆45Jul 13, 2026Updated 2 months ago
- Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster☆76Jul 25, 2026Updated 2 months ago
- Run on TWO-DGX-Spark - vLLm-0.24.0 dual cache optimized DSV4F+DSpark+NVFP4 KV (Concurrency 12 with 1.5M context/3M KV token Pool) >0.58-0…☆21Aug 17, 2026Updated last month
- MiMo-V2.5 Omni TP=2 on 2x DGX Spark · 1M context · NVFP4 4-bit KV (~1.97M-token KV pool @ 1M, ~30 tok/s) · 69-eval: thinking-OFF 97.8 bea…☆39Jul 13, 2026Updated 2 months ago
- Give any blind LLM eyes — on any machine — in one shot. Tiny local VLM (Qwen3.5-0.8B) as an OpenAI-compatible vision service. Mac (MLX) /…☆28Jul 7, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Run the AEON Bench suite on your own hardware: verified HuggingFace pull → serve → benchmark (text · agentic ×3 harnesses · vision · audi…☆27Sep 11, 2026Updated 3 weeks ago
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆107Aug 20, 2026Updated last month
- Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker…☆130Updated this week
- Mixed-capability LLM benchmark for DGX Spark — 57 scenarios, 10 domains, partial-credit grading, trial statistics☆164Aug 29, 2026Updated last month
- Open Source Desktop App to help with Previs for AI-native filmmaking — stage grey-box scenes, choreograph camera & cast with marks, expor…☆160Aug 5, 2026Updated last month
- ☆93Updated this week
- ☆19Mar 21, 2026Updated 6 months ago
- Watch-sized 4K USB-C IP-KVM☆41Sep 4, 2026Updated 3 weeks ago
- sparkDash ⚡ — Multi-DGX Spark Monitoring Dashboard☆515Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Homebridge plugin for Pentair IntelliCenter☆12Apr 13, 2023Updated 3 years ago
- my version of the models dev site which is way faster and has the filters I want☆18Apr 22, 2026Updated 5 months ago
- HiddenVM is a futuristic tool powered by KVM designed to combine the powerful amnesic nature of Tails and the impenetrable design of Whon…☆11Jul 22, 2022Updated 4 years ago
- A language model that forms persistent memories from conversation and maintains them through sleep. MEMIT weight editing + null-space-con…☆75Jul 11, 2026Updated 2 months ago
- ☆16Sep 15, 2026Updated 2 weeks ago
- A self-evolving agent in rust with one tool that only uses skills☆31Aug 12, 2026Updated last month
- Swift package for reading and writing Safetensors files.☆13Feb 6, 2026Updated 7 months ago
- Control panel for VLLM, Sglang, llama.cpp, exllamav3☆1,811Updated this week
- Serve the home! Inference stack for your Nvidia DGX Spark aka the Grace Blackwell AI supercomputer on your desk. Mostly vLLM based for no…☆52Sep 20, 2026Updated last week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆17Sep 5, 2021Updated 5 years ago
- Qwen3.6-35B-A3B-heretic NVFP4 + DFlash speculative decoding on DGX Spark (GB10/sm_121a). Source-built vLLM image + 7 patches + comprehens…☆144Jun 28, 2026Updated 3 months ago
- Sensirion SVM30 on Arduino, ESPxx, 32U4, Lorawan, DUE☆10Oct 30, 2023Updated 2 years ago
- ☆15Feb 3, 2026Updated 8 months ago
- DGX Spark / GB10 vLLM image for Gemma 4 31B Deckard Heretic Uncensored NVFP4 with z-lab DFlash speculative decoding.☆34May 15, 2026Updated 4 months ago
- ☆16May 21, 2026Updated 4 months ago
- Local benchmarking UI for LLMs and AI agents☆24Apr 13, 2026Updated 5 months ago
- AseeVR fix for Droolon Pi1☆14Mar 16, 2024Updated 2 years ago
- LibreTuya ESPHome Home Assistant Add-on☆18Apr 17, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- If I had my druthers....a programming language made during Khan Academy's 2014 Hack Week.☆15Jan 26, 2015Updated 11 years ago
- CPU-side action routing, context compression, and causal memory for AI agents — matches an LLM-everything agent at 58% fewer LLM calls an…☆39Updated this week
- sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems. Live support: https://discord.com/invite/GH5kRg…☆526Updated this week
- Project Doxa is an advanced, autonomous agent-based simulation engine where AI agents act as citizens of evolving civilizations.☆27Jun 14, 2026Updated 3 months ago
- A ComfyUI custom node that brings a DAW-style interactive video timeline directly into the node graph. Upload any video, scrub through it…☆19Apr 21, 2026Updated 5 months ago
- Swift Package Manager package for embedding Python in Swift app☆21Feb 15, 2024Updated 2 years ago
- ☆23Mar 10, 2022Updated 4 years ago