DeepSeek V4 Flash specific inference engine. SSD MoE expert paging (slot-bank) + disk KV cache for long agent sessions. Metal-first, narrow & high-performance.
☆32Sep 6, 2026Updated 3 weeks ago
Alternatives and similar repositories for ds4-ssd
Users that are interested in ds4-ssd are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆38Mar 30, 2026Updated 6 months ago
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆38Jun 12, 2026Updated 3 months ago
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆133Sep 4, 2026Updated 3 weeks ago
- NetHack as-a-library. Semantic world exploration engine.☆52Sep 14, 2026Updated 2 weeks ago
- this repo has all official MLX-LM-LoRA example notebooks for training on Apple Silicon☆42Sep 13, 2026Updated 2 weeks ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Instant Perfect Native MacOS Transcription☆55Jul 26, 2025Updated last year
- An MLX port of Meta's Coconut reasoning model☆16Sep 2, 2025Updated last year
- import documents for LLMs☆49Updated this week
- ☆21Oct 9, 2024Updated last year
- General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.☆17Apr 26, 2026Updated 5 months ago
- ☆43Jun 15, 2026Updated 3 months ago
- My reasearch of losslessly compressing LLM weights.☆65Jul 23, 2026Updated 2 months ago
- Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs — speed, quality, live dashboard, publisha…☆40Updated this week
- RLM (Recursive Language Model) extension for pi - process large context files that exceed LLM context windows☆19Feb 8, 2026Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆16Feb 21, 2026Updated 7 months ago
- revm (Rust Ethereum VM) translation for Era / zkEVM☆13Jun 20, 2026Updated 3 months ago
- Zero-Knowledge Proof for Nuclear Disarmament Verification☆15Dec 7, 2023Updated 2 years ago
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆107Aug 20, 2026Updated last month
- ☆25Jun 30, 2026Updated 3 months ago
- ☆17Aug 13, 2026Updated last month
- a SplineCamera react component☆14Feb 18, 2024Updated 2 years ago
- Generate zero-knowledge proofs of valid RSA signatures from your browser.☆15Feb 23, 2023Updated 3 years ago
- Local, model-agnostic image-generation studio for Apple silicon (MLX). Clean Claude-design UI; ships with the Krea 2 Turbo pure-MLX backe…☆45Aug 4, 2026Updated last month
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Tools for merging pretrained large language models.☆19Jun 12, 2024Updated 2 years ago
- ☆32Feb 2, 2025Updated last year
- TRAMP-like transparent remote execution for pi — tools run remotely via SSH/Docker, pi stays local☆20May 23, 2026Updated 4 months ago
- ☆18May 14, 2024Updated 2 years ago
- ☆19Jun 18, 2024Updated 2 years ago
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 6 months ago
- Daily proof-of-concept implementations of cutting-edge AI advancements☆30Mar 20, 2026Updated 6 months ago
- bf16 LoRA fine-tuning of [Qwen3.5-35B-A3B](https://huggingface.co/unsloth/Qwen3.5-35B-A3B) (a 35B-total / 3B-active Mixture-of-Experts vi…☆18Mar 12, 2026Updated 6 months ago
- ☆10Jul 14, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Codebase for VidHal: Benchmarking Hallucinations in Vision LLMs☆14Apr 23, 2026Updated 5 months ago
- 📊 LLM Context Benchmarks - A comprehensive benchmarking tool for testing LLMs with varying context sizes using Ollama. Features dual b…☆99Updated this week
- A library to automate the conversion of linux-based VMs to a set of docker containers☆14Apr 10, 2015Updated 11 years ago
- ansible playbook to create HA kubernetes cluster☆10Jul 25, 2018Updated 8 years ago
- Run controlnet with flux☆17Oct 8, 2024Updated last year
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆61Apr 18, 2026Updated 5 months ago
- Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — …☆83Jul 13, 2026Updated 2 months ago