DeepSeek V4 Flash specific inference engine. SSD MoE expert paging (slot-bank) + disk KV cache for long agent sessions. Metal-first, narrow & high-performance.
☆31Jul 18, 2026Updated last month
Alternatives and similar repositories for ds4-ssd
Users that are interested in ds4-ssd are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆38Jun 12, 2026Updated 2 months ago
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆129Aug 15, 2026Updated last week
- this repo has all official MLX-LM-LoRA example notebooks for training on Apple Silicon☆38Apr 23, 2026Updated 4 months ago
- Instant Perfect Native MacOS Transcription☆54Jul 26, 2025Updated last year
- An MLX port of Meta's Coconut reasoning model☆16Sep 2, 2025Updated 11 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Train Embedding Models on MLX.☆17Jun 2, 2026Updated 2 months ago
- import documents for LLMs☆48Jul 7, 2026Updated last month
- General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.☆17Apr 26, 2026Updated 3 months ago
- ☆21Oct 9, 2024Updated last year
- My reasearch of losslessly compressing LLM weights.☆63Jul 23, 2026Updated last month
- RLM (Recursive Language Model) extension for pi - process large context files that exceed LLM context windows☆18Feb 8, 2026Updated 6 months ago
- ☆16Feb 21, 2026Updated 6 months ago
- revm (Rust Ethereum VM) translation for Era / zkEVM☆13Jun 20, 2026Updated 2 months ago
- Gradio chat interface for FastMLX☆12Sep 22, 2024Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Recibo: encrypted memos for ERC-20 transactions.☆15Jun 18, 2026Updated 2 months ago
- Local, model-agnostic image-generation studio for Apple silicon (MLX). Clean Claude-design UI; ships with the Krea 2 Turbo pure-MLX backe…☆43Aug 4, 2026Updated 2 weeks ago
- Gonçalo's instructions and skills for agents☆24Jul 31, 2026Updated 3 weeks ago
- a SplineCamera react component☆14Feb 18, 2024Updated 2 years ago
- Let anyone submit a coding task against your codebase.☆21Feb 13, 2026Updated 6 months ago
- An iOS app to allow you to verify a picture was taken on an iPhone via hardware signing and cryptographic proving. Inspired as the revers…☆13Feb 25, 2024Updated 2 years ago
- TRAMP-like transparent remote execution for pi — tools run remotely via SSH/Docker, pi stays local☆20May 23, 2026Updated 3 months ago
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 5 months ago
- Links to all the source code and solutions I reference in my O'Reilly Introduction to Docker video tutorial☆11Dec 10, 2014Updated 11 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆51Jul 30, 2025Updated last year
- skill.md specialized skill registry for AI agents☆15Apr 24, 2026Updated 3 months ago
- ☆17May 8, 2024Updated 2 years ago
- 📊 LLM Context Benchmarks - A comprehensive benchmarking tool for testing LLMs with varying context sizes using Ollama. Features dual b…☆81Updated this week
- Run controlnet with flux☆17Oct 8, 2024Updated last year
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆61Apr 18, 2026Updated 4 months ago
- A Prometheus metrics exporter for NVIDIA DGX Spark clusters.☆21Feb 16, 2026Updated 6 months ago
- MLX Implementation of Recursive Reasoning with Tiny Networks☆79Oct 11, 2025Updated 10 months ago
- This repo maintains a 'cheat sheet' for LLMs that are undertrained on mlx☆33Mar 12, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆28May 13, 2026Updated 3 months ago
- ANE (Apple Neural Engine) CostModel profiler for CoreML models☆36Apr 9, 2026Updated 4 months ago
- ☆47Sep 30, 2025Updated 10 months ago
- ☆31Apr 22, 2026Updated 4 months ago
- ☆11Sep 11, 2024Updated last year
- Community model zoo for Apple Core AI (iOS/macOS 27): 62 models — LLM, VLM, OCR, ASR, TTS, image/video/music gen, forecasting — each gate…☆394Updated this week
- MLX version of DINO DETR☆17Dec 26, 2024Updated last year