vLLM 0.25.1 serving stack for poolside/Laguna-S-2.1-NVFP4 with DFlash speculative decoding — DGX Spark & RTX 6000 PRO
☆54Jul 22, 2026Updated this week
Alternatives and similar repositories for Laguna-S-2.1-DGX-Spark-RTX-6000-PRO
Users that are interested in Laguna-S-2.1-DGX-Spark-RTX-6000-PRO are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆84Jul 9, 2026Updated 2 weeks ago
- sparkDash ⚡ — Multi-DGX Spark Monitoring Dashboard☆82Updated this week
- ☆166Updated this week
- DeepSeek-V4-Flash-DSpark abliterated (uncensored) · ~100% refusal bypass · C1 ~57 tok/s · 1M ctx · 2× DGX Spark · HF weights☆22Jul 12, 2026Updated last week
- MiniMax-M3 (428B, no pruning) at 36 tok/s on 2× NVIDIA DGX Spark — W4A16 GPTQ + NVFP4 KV + EAGLE-3 speculative decoding on vLLM. Three se…☆39Jul 13, 2026Updated last week
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Tencent Hunyuan 3 (295B MoE) on 2x NVIDIA DGX Spark: NVFP4 W4A16 + native MTP speculative decoding. First published MTP-on-GB10 numbers, …☆18Jul 13, 2026Updated last week
- MiMo-V2.5 Omni TP=2 on 2x DGX Spark · 1M context · NVFP4 4-bit KV (~1.97M-token KV pool @ 1M, ~30 tok/s) · 69-eval: thinking-OFF 97.8 bea…☆37Jul 13, 2026Updated last week
- Live dashboard: fire N parallel streaming coding-agent runs at any OpenAI-compatible endpoint — per-run TTFT/tok-s/E2E, real kill switch,…☆20Jul 8, 2026Updated 2 weeks ago
- llama-server start/stop scripts for Qwen3.6-35B-A3B UD-Q8_K_XL GGUF on DGX Spark☆24Jul 6, 2026Updated 2 weeks ago
- NVFP4 Gemma-4 26B-A4B MoE for DGX Spark — optimal recipe: DFlash n=10 (flex) on AEON vLLM Ultimate. 144 tok/s single / 1,724 peak (Coding…☆42Jun 28, 2026Updated 3 weeks ago
- AEON vLLM Ultimate — vLLM 0.25.0 built from source for DGX Spark / Blackwell (sm_121a/GB10). One image serves the whole AEON fleet (Gemma…☆98Jul 17, 2026Updated last week
- Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker…☆92Jul 11, 2026Updated last week
- Background thinking plugin for Hermes Agent — MCTS-powered autonomous dream system☆20May 5, 2026Updated 2 months ago
- DeepSeek-V4-Flash on a Raspberry Pi 5 (8GB)☆21Jun 9, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆13Jun 20, 2022Updated 4 years ago
- Divisive Intelligent K-Means algorithm (DiviK) for joint feature selection and clustering of heavily multidimensional data.☆15Jul 9, 2024Updated 2 years ago
- Profile repo — categorized index of NVFP4 model releases, DGX Spark inference stacks, Apple Silicon MLX builds, the voice-AI stack, and t…☆37Jul 9, 2026Updated 2 weeks ago
- Code for "HyperTab: Hypernetwork Approach for Deep Learning on Small Tabular Datasets"☆16Feb 12, 2024Updated 2 years ago
- Axway ATS Test Explorer - Web application for exploring test results☆14Oct 3, 2024Updated last year
- Bleeding-edge ComfyUI for NVIDIA DGX Spark (GB10/Blackwell/sm_121a). CUDA 13 + SageAttention v3 (sm_121a) + NVFP4 + 14 custom-node packs …☆71Jun 28, 2026Updated 3 weeks ago
- Run the AEON Bench suite on your own hardware: verified HuggingFace pull → serve → benchmark (text · agentic ×3 harnesses · vision · audi…☆21Updated this week
- Mixed-capability LLM benchmark for DGX Spark — 57 scenarios, 10 domains, partial-credit grading, trial statistics☆136Jul 11, 2026Updated last week
- A copy of my Mathematics and Computer Engineering B.Sc. thesis☆19Dec 8, 2020Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- yara and radare2, better together☆28Apr 13, 2026Updated 3 months ago
- Retypd plugin for Ghidra reverse engineering framework from NSA☆28Jul 6, 2023Updated 3 years ago
- Source code for XTRIDE: "Practical Type Inference: High-Throughput Recovery of Real-World Structures and Function Signatures"☆26Jun 24, 2026Updated last month
- Open Source Desktop App for Local-first pre-production canvas for AI video planning, prompts, assets, and handoff packages.☆32Updated this week
- ☆24Feb 18, 2025Updated last year
- PDB Rewriting Rust Library☆29Apr 26, 2024Updated 2 years ago
- Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — …☆57Jul 13, 2026Updated last week
- UI for Pandas AI, the Python library that makes dataframes conversational.☆14Jun 6, 2023Updated 3 years ago
- Collection of step-by-step playbooks for setting up AI/ML workloads on NVIDIA DGX Spark devices with Blackwell architecture.☆1,172Updated this week
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Local benchmarking UI for LLMs and AI agents☆20Apr 13, 2026Updated 3 months ago
- Simple pyethapp development environment☆30Apr 27, 2018Updated 8 years ago
- ☆21Updated this week
- ☆23Oct 18, 2021Updated 4 years ago
- A relational logic programming language embedded in Rust.☆12Aug 15, 2025Updated 11 months ago
- Example on how to use a Tensorflow Queue to feed data to your models.☆38Mar 30, 2020Updated 6 years ago
- Better type id and Any for Rust☆17Dec 12, 2025Updated 7 months ago