DeepSeek V4 Flash specific inference engine. SSD MoE expert paging (slot-bank) + disk KV cache for long agent sessions. Metal-first, narrow & high-performance.
☆29Jul 18, 2026Updated 2 weeks ago
Alternatives and similar repositories for ds4-ssd
Users that are interested in ds4-ssd are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆37Jun 12, 2026Updated last month
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆117Jul 15, 2026Updated 2 weeks ago
- Instant Perfect Native MacOS Transcription☆54Jul 26, 2025Updated last year
- An MLX port of Meta's Coconut reasoning model☆16Sep 2, 2025Updated 11 months ago
- Train Embedding Models on MLX.☆17Jun 2, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- import documents for LLMs☆48Jul 7, 2026Updated 3 weeks ago
- General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.☆17Apr 26, 2026Updated 3 months ago
- Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs — speed, quality, live dashboard, publisha…☆24Jul 19, 2026Updated 2 weeks ago
- ☆21Oct 9, 2024Updated last year
- ☆43Jun 15, 2026Updated last month
- My reasearch of losslessly compressing LLM weights.☆62Jul 23, 2026Updated last week
- Universal Wallet for Ethereum☆10Jan 13, 2025Updated last year
- RLM (Recursive Language Model) extension for pi - process large context files that exceed LLM context windows☆15Feb 8, 2026Updated 5 months ago
- ☆16Feb 21, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Gradio chat interface for FastMLX☆12Sep 22, 2024Updated last year
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆91Jul 9, 2026Updated 3 weeks ago
- Local, model-agnostic image-generation studio for Apple silicon (MLX). Clean Claude-design UI; ships with the Krea 2 Turbo pure-MLX backe…☆43Jul 15, 2026Updated 2 weeks ago
- Gonçalo's instructions and skills for agents☆22Updated this week
- ☆24Jun 30, 2026Updated last month
- ☆32Feb 2, 2025Updated last year
- TRAMP-like transparent remote execution for pi — tools run remotely via SSH/Docker, pi stays local☆20May 23, 2026Updated 2 months ago
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 4 months ago
- Daily proof-of-concept implementations of cutting-edge AI advancements☆30Mar 20, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Links to all the source code and solutions I reference in my O'Reilly Introduction to Docker video tutorial☆11Dec 10, 2014Updated 11 years ago
- ☆10Jul 14, 2025Updated last year
- 마크다운으로 한글 소설을 쓰기 위한 boilerplate☆20Feb 2, 2022Updated 4 years ago
- Codebase for VidHal: Benchmarking Hallucinations in Vision LLMs☆14Apr 23, 2026Updated 3 months ago
- ☆17May 8, 2024Updated 2 years ago
- 📊 LLM Context Benchmarks - A comprehensive benchmarking tool for testing LLMs with varying context sizes using Ollama. Features dual b…☆78Updated this week
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆62Apr 18, 2026Updated 3 months ago
- Run controlnet with flux☆17Oct 8, 2024Updated last year
- Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — …☆63Jul 13, 2026Updated 3 weeks ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Open-source Claude skills and an MCP server for Sales and GTM teams: build Sales Navigator searches, write multichannel campaigns, and ha…☆30Jul 24, 2026Updated last week
- MLX Implementation of Recursive Reasoning with Tiny Networks☆79Oct 11, 2025Updated 9 months ago
- This repo maintains a 'cheat sheet' for LLMs that are undertrained on mlx☆33Mar 12, 2026Updated 4 months ago
- ☆28May 13, 2026Updated 2 months ago
- vibevoice real time 0.5B swift port☆31Dec 12, 2025Updated 7 months ago
- ANE (Apple Neural Engine) CostModel profiler for CoreML models☆36Apr 9, 2026Updated 3 months ago
- ☆45Sep 30, 2025Updated 10 months ago