Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpark, and full continuous batch support
☆405Aug 27, 2026Updated last month
Alternatives and similar repositories for ds4-on-spark
Users that are interested in ds4-on-spark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16Jun 27, 2026Updated 3 months ago
- Qwen3.5-122B-A10B on a DGX Spark with DFlash speculative decode. One-shot Docker/vLLM installer. 80+ tok/s!☆59Jun 29, 2026Updated 3 months ago
- Some benchmark results of small models and quants that fit on DGX Spark☆51Aug 23, 2026Updated last month
- Qwen3.6-35B-A3B-heretic NVFP4 + DFlash speculative decoding on DGX Spark (GB10/sm_121a). Source-built vLLM image + 7 patches + comprehens…☆144Jun 28, 2026Updated 3 months ago
- A monitor of resources for DGX Spark☆25Feb 13, 2026Updated 7 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)☆315Aug 23, 2026Updated last month
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆102Updated this week
- Operator-grade GPU monitor for NVIDIA GPUs with native GB10 / DGX Spark coherent UMA support — PSI pressure, clock detection, ConnectX-7 …☆33May 31, 2026Updated 4 months ago
- DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark☆490Sep 10, 2026Updated 3 weeks ago
- Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.☆128Sep 10, 2026Updated 3 weeks ago
- Web UI for sparkrun — launch and monitor inference workloads on NVIDIA DGX Spark☆28Jun 16, 2026Updated 3 months ago
- sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems. Live support: https://discord.com/invite/GH5kRg…☆526Updated this week
- 🚀 Multi-agent orchestration for Qwen Code CLI. 24 specialized AI agents and 82+ skills working as a professional dev department. Built f…☆50Jun 28, 2026Updated 3 months ago
- Official Spark Arena Recipe Registry☆64Sep 3, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.i…☆469Sep 15, 2026Updated 2 weeks ago
- Self-hosted voice for coding agents. Talk from any browser or a Telegram call, interrupt mid-sentence, clone any voice, and hand real wor…☆41Updated this week
- FlashQLA TileLang GDN kernels ported to NVIDIA Blackwell consumer (GB10 / DGX Spark)☆18Jun 5, 2026Updated 3 months ago
- Collection of step-by-step playbooks for setting up AI/ML workloads on NVIDIA DGX Spark devices with Blackwell architecture.☆1,397Updated this week
- ☆21Apr 7, 2026Updated 5 months ago
- Zero-knowledge encrypted pastebin. Your data is encrypted before it leaves your browser.☆35Sep 5, 2026Updated 3 weeks ago
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆44Jun 28, 2026Updated 3 months ago
- NEURA Office: native Word, PowerPoint and Excel for Open WebUI. LibreOffice-compatible OOXML. MS365 assistant planned.☆74Sep 17, 2026Updated 2 weeks ago
- Unify every AI agent on ONE local shared memory over MCP. Obsidian vault + gbrain (by Garry Tan) + Ollama. 100% offline, $0 API. Orchestr…☆51Jul 25, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- DGX Spark inference dashboard — vLLM, SGLang, llama.cpp, WebGPU & sparkrun with agent-powered Auto-Fix and speed optimization☆66Updated this week
- ☆14Updated this week
- LLM proxy specialized in analyzing harness traffic to catch cache invalidation.☆58Updated this week
- Self-improving LLM system using Generator-Reflector-Curator pattern for online learning from execution feedback☆36Aug 11, 2026Updated last month
- LLM Stack for nVidia DGX Spark containing LiteLLM, LamaSwap, vLLM, Llama.cpp and ollama☆46Sep 22, 2026Updated last week
- Rust implementation of TurboQuant, PolarQuant, and QJL — zero-overhead vector quantization for semantic search and KV cache compression (…☆26Updated this week
- Index your shell history in a full-text search database☆13Aug 23, 2024Updated 2 years ago
- Coding agent with cool features.☆19Aug 23, 2026Updated last month
- SAM 3D bodyを用いて画像から3D人形(Blender用のメッシュとボーン▶FBX、ClipStudioPaint用のポーズデータ▶BVH)を取り出すツールです☆22Jan 28, 2026Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A lightweight orchestration framework that piggybacks your local Agentic CLI setup. Intentionally simple. Yet powerful... like an army of…☆26Jun 21, 2026Updated 3 months ago
- DeepSeek-V4.1-Flash (552B MoE, MXFP4 experts, 1M ctx) on four NVIDIA DGX Sparks with vLLM TP4: Engram-on-disk patch, sm121 kernel build, …☆89Sep 19, 2026Updated 2 weeks ago
- AI Context Takt: A pipeline tool for structural control over LLM context. Escape the black box of chat history, maximize token efficiency…☆21Jan 12, 2026Updated 8 months ago
- Multi-vector latent space steering adapter module for language models☆20Nov 22, 2025Updated 10 months ago
- Deploying full-stack on-prem deep research agent that can be run entirely on a local machine for $0!☆34Nov 8, 2025Updated 10 months ago
- From-scratch voice agents in Python: end-to-end speech pipelines, runnable chapters, and a small shared library. Local models, explicit s…☆47May 3, 2026Updated 5 months ago
- Bleeding-edge ComfyUI for NVIDIA DGX Spark (GB10/Blackwell/sm_121a). CUDA 13 + SageAttention v3 (sm_121a) + NVFP4 + 14 custom-node packs …☆86Jun 28, 2026Updated 3 months ago