DeepSeek V4 Flash specific inference engine. SSD MoE expert paging (slot-bank) + disk KV cache for long agent sessions. Metal-first, narrow & high-performance.
☆31Sep 6, 2026Updated this week
Alternatives and similar repositories for ds4-ssd
Users that are interested in ds4-ssd are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆37Mar 30, 2026Updated 5 months ago
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆38Jun 12, 2026Updated 3 months ago
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆132Sep 4, 2026Updated last week
- NetHack as-a-library. Semantic world exploration engine.☆49Updated this week
- this repo has all official MLX-LM-LoRA example notebooks for training on Apple Silicon☆41Apr 23, 2026Updated 4 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Instant Perfect Native MacOS Transcription☆54Jul 26, 2025Updated last year
- An MLX port of Meta's Coconut reasoning model☆16Sep 2, 2025Updated last year
- Train Embedding Models on MLX.☆17Jun 2, 2026Updated 3 months ago
- import documents for LLMs☆49Jul 7, 2026Updated 2 months ago
- ☆21Oct 9, 2024Updated last year
- General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.☆17Apr 26, 2026Updated 4 months ago
- ☆43Jun 15, 2026Updated 2 months ago
- Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs — speed, quality, live dashboard, publisha…☆37Updated this week
- RLM (Recursive Language Model) extension for pi - process large context files that exceed LLM context windows☆18Feb 8, 2026Updated 7 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆16Feb 21, 2026Updated 6 months ago
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆106Aug 20, 2026Updated 3 weeks ago
- Gonçalo's instructions and skills for agents☆24Updated this week
- ☆25Jun 30, 2026Updated 2 months ago
- a SplineCamera react component☆14Feb 18, 2024Updated 2 years ago
- [ICML 2026] Scaling Beyond Masked Diffusion Language Models☆37Updated this week
- Tools for merging pretrained large language models.☆19Jun 12, 2024Updated 2 years ago
- Claude Code's compaction engine, as a drop-in DeepAgents middleware☆46Aug 26, 2026Updated 2 weeks ago
- Let anyone submit a coding task against your codebase.☆21Feb 13, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- TRAMP-like transparent remote execution for pi — tools run remotely via SSH/Docker, pi stays local☆20May 23, 2026Updated 3 months ago
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 5 months ago
- Daily proof-of-concept implementations of cutting-edge AI advancements☆30Mar 20, 2026Updated 5 months ago
- bf16 LoRA fine-tuning of [Qwen3.5-35B-A3B](https://huggingface.co/unsloth/Qwen3.5-35B-A3B) (a 35B-total / 3B-active Mixture-of-Experts vi…☆18Mar 12, 2026Updated 6 months ago
- Links to all the source code and solutions I reference in my O'Reilly Introduction to Docker video tutorial☆11Dec 10, 2014Updated 11 years ago
- ☆10Jul 14, 2025Updated last year
- skill.md specialized skill registry for AI agents☆15Apr 24, 2026Updated 4 months ago
- ☆17May 8, 2024Updated 2 years ago
- 📊 LLM Context Benchmarks - A comprehensive benchmarking tool for testing LLMs with varying context sizes using Ollama. Features dual b…☆91Aug 30, 2026Updated 2 weeks ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A library to automate the conversion of linux-based VMs to a set of docker containers☆14Apr 10, 2015Updated 11 years ago
- ansible playbook to create HA kubernetes cluster☆10Jul 25, 2018Updated 8 years ago
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆61Apr 18, 2026Updated 4 months ago
- Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — …☆82Jul 13, 2026Updated 2 months ago
- A Prometheus metrics exporter for NVIDIA DGX Spark clusters.☆23Feb 16, 2026Updated 6 months ago
- Open-source Claude skills and an MCP server for Sales and GTM teams: build Sales Navigator searches, write multichannel campaigns, and ha…☆37Updated this week
- MLX Implementation of Recursive Reasoning with Tiny Networks☆79Oct 11, 2025Updated 11 months ago