☆32Jul 17, 2026Updated last week
Alternatives and similar repositories for GLM5.2-2bit-2-DGX-Spark--21.5tok-s
Users that are interested in GLM5.2-2bit-2-DGX-Spark--21.5tok-s are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DeepSeek-V4-Flash-DSpark abliterated (uncensored) · ~100% refusal bypass · C1 ~57 tok/s · 1M ctx · 2× DGX Spark · HF weights☆22Jul 12, 2026Updated 2 weeks ago
- MiniMax-M3 (428B, no pruning) at 36 tok/s on 2× NVIDIA DGX Spark — W4A16 GPTQ + NVFP4 KV + EAGLE-3 speculative decoding on vLLM. Three se…☆39Jul 13, 2026Updated 2 weeks ago
- MiMo-V2.5 Omni TP=2 on 2x DGX Spark · 1M context · NVFP4 4-bit KV (~1.97M-token KV pool @ 1M, ~30 tok/s) · 69-eval: thinking-OFF 97.8 bea…☆37Jul 13, 2026Updated 2 weeks ago
- Run on TWO-DGX-Spark - vLLm-0.24.0 dual cache optimized DSV4F+DSpark+NVFP4 KV (Concurrency 12 with 1.5M context/3M KV token Pool) >0.58-0…☆20Jul 7, 2026Updated 2 weeks ago
- A PyTorch native library for large model training☆29Apr 1, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster☆70Updated this week
- Self-contained DeepSeek V4 Flash DSpark TP=2 recipe for 2x DGX Spark with 62 tok/s benchmark☆22Jul 13, 2026Updated 2 weeks ago
- Examples for using the SiLLM framework for training and running Large Language Models (LLMs) on Apple Silicon☆16May 8, 2025Updated last year
- vLLM 0.25.1 serving stack for poolside/Laguna-S-2.1-NVFP4 with DFlash speculative decoding — DGX Spark & RTX 6000 PRO☆57Updated this week
- ☆12May 28, 2025Updated last year
- An open-source hardware racing drone design and demonstration RL software for the artificial intelligence robotic drone race competitions☆12Sep 6, 2020Updated 5 years ago
- Forgotten Developer - a terminal like theme for Frontity.☆11Jun 6, 2022Updated 4 years ago
- Langchain + Docker + Neo4j☆10Oct 29, 2024Updated last year
- Learning Pytorch☆13Jun 12, 2018Updated 8 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- The world's smallest AI agent. ESP32 + pure C. $4 chip, ~120KB RAM.☆15Mar 4, 2026Updated 4 months ago
- vLLM + Qwen3.5-122B-A10B-NVFP4 on NVIDIA DGX Spark (GB10/SM121) — single-GPU NVFP4 W4A4 with MTP speculative decoding, self-contained Doc…☆40Mar 12, 2026Updated 4 months ago
- Ollama Modelfiles - Discover more at OllamaHub☆20Dec 2, 2023Updated 2 years ago
- Light weight Object detection on Nintendo 3DS, powered by NCNN☆13Apr 3, 2024Updated 2 years ago
- Quickly inference EdgeSAM with MNN☆15Dec 19, 2023Updated 2 years ago
- Production-ready vLLM deployment wrapper for Qwen3.6-27B (NVFP4) — self-hosted OpenAI-compatible inference☆47Jul 1, 2026Updated 3 weeks ago
- O **vOx Oratória** é uma aplicação web **Open Source** e **Local-First**, projetada para democratizar o acesso ao treinamento de comunica…☆16Dec 1, 2025Updated 7 months ago
- A Claude Code plugin for a structured dev workflow — 14 expert agents, code review, planning, and task tracking, all from the terminal.☆16Updated this week
- ☆14Sep 30, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Repository used to main group ACLs used by Kubeflow developers☆19Updated this week
- A Swift library for making Alfred Workflows☆21Jan 31, 2017Updated 9 years ago
- Distort a video using Seam Carving (video) and Vibrato effect (sound)☆16Nov 3, 2021Updated 4 years ago
- A simplified implementation of RetinaNet from https://arxiv.org/pdf/1708.02002.pdf using TF2.0☆13Aug 5, 2020Updated 5 years ago
- This is an LLM interface that you can use to analyze and get insight into diary entries or other documents completely offline.☆16Dec 31, 2023Updated 2 years ago
- A small example repo for building a backend with LlamaParse☆21May 13, 2026Updated 2 months ago
- Slides and code for PyOhio 2017 presentation☆17Jul 30, 2017Updated 8 years ago
- Cloud Observability Gemini CLI Extension☆28Sep 23, 2025Updated 10 months ago
- DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark☆157Jul 16, 2026Updated last week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A tool to assist in the interpretation of learned features in sparse autoencoders (in particular the four SAE's trained by Joseph Bloom o…☆19Oct 4, 2024Updated last year
- VXPromotionViewController is a simple inline and cross promotion display for your iOS 7 app. It can load the app information from the App…☆24Oct 17, 2019Updated 6 years ago
- Dual-Branch Meta-learning Network with Distribution Alignment for Face Anti-spoofing, Transactions on Information Forensics & Security☆17Jan 5, 2022Updated 4 years ago
- {{moment}} handlebars helper. Combines the powers of Assemble, Handlebars.js and Moment.js into a great helper to master time.☆19Jan 3, 2018Updated 8 years ago
- The official implementation of Preference Data Reward-Augmentation.☆18May 1, 2025Updated last year
- Spin up a full medusa store with docker.☆32Feb 21, 2023Updated 3 years ago
- "yield" for Swift, inspired by Python and F#☆23Apr 1, 2015Updated 11 years ago