3.34× faster inference on Apple Silicon: native MLX port of DFlash speculative decoding
☆19Oct 7, 2026Updated this week
Alternatives and similar repositories for mlx-dflash
Users that are interested in mlx-dflash are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆18Dec 1, 2025Updated 10 months ago
- DocFinder is a local-first indexing and searching documents using semantic embeddings stored in SQLite. Everything runs on your machine, …☆33Oct 1, 2026Updated last week
- A Prometheus metrics exporter for NVIDIA DGX Spark clusters.☆25Feb 16, 2026Updated 7 months ago
- ☆20Jan 3, 2026Updated 9 months ago
- A Chrome extension that enables virtual fashion try-on and model swap using FASHN AI. Hover over fashion images on any website to: (1) tr…☆22Aug 14, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Inspect LLM's logprobs and perplexity over a piece of text, or compare two LLMs (like a git diff)☆18Aug 3, 2026Updated 2 months ago
- A robust Python toolkit for converting video/audio content into accurate, multilingual subtitles using WhisperX for transcription and Goo…☆29Sep 16, 2026Updated 3 weeks ago
- ☆42Jul 13, 2026Updated 2 months ago
- A MCP stdio toolpack for local LLMs☆34Apr 6, 2026Updated 6 months ago
- An educational Rust project for exporting and running inference on Qwen3 LLM family☆44Aug 3, 2025Updated last year
- A curated collection of persona-based mcp server & tool groupings.☆39Sep 11, 2025Updated last year
- A miniaturized version of the Kimi-K2 model optimized for deployment on single H100 GPUs.☆36Jul 16, 2025Updated last year
- Nats jetstream queue adapter with client for Laravel framework☆36Jun 29, 2026Updated 3 months ago
- open-source, local-first AI CLI for developers. It connects to your local Ollama models or remote providers like OpenAI, Anthropic, Gemin…☆61Jul 2, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆24Apr 3, 2026Updated 6 months ago
- Speculative Decoding Implementations: MTP, EAGLE-3, Medusa-1, PARD, Draft Models, N-gram and Suffix Decoding from scratch☆15May 2, 2026Updated 5 months ago
- npm package template with typescript and tsup☆11Nov 27, 2025Updated 10 months ago
- Train Llama 3 models from scratch. Any scale, any personality. By Arianna Method.☆52May 4, 2026Updated 5 months ago
- For the better CI as well as CD using gogs and drone base on kubernetes☆10Jul 31, 2021Updated 5 years ago
- Tiny, composable Atomic CSS engine☆13Apr 1, 2022Updated 4 years ago
- Load and run Llama from safetensors files in C☆16Oct 24, 2024Updated last year
- ☆11Feb 28, 2022Updated 4 years ago
- Vite utility for vue3 server side rendering☆10Jul 21, 2026Updated 2 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Composition of Multimodal Language Models From Scratch☆16Aug 16, 2024Updated 2 years ago
- A hands-on guide for AI builders: make your own RTX 4090D/5090 GPU server that’s fast and efficient.☆17Aug 19, 2025Updated last year
- Options pricing, Greeks, strategy P&L, volatility surfaces, and scenario analysis.☆15Sep 24, 2026Updated 2 weeks ago
- Desktop app with Compose Multiplatform to use Qwen3-TTS with an UI.☆78Aug 2, 2026Updated 2 months ago
- Code and Data for Evaluating the Evaluators☆16Aug 20, 2025Updated last year
- ☆16Feb 1, 2025Updated last year
- General Tool-calling API Proxy☆61Mar 26, 2026Updated 6 months ago
- ☆11May 2, 2023Updated 3 years ago
- SvelteKit GraphQL queries using fetch only: how you can drop Apollo client and urql dependencies altogether to make your Svelte app leane…☆16Jul 23, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 🎨 TailwindCSS utility to override background fill color using shadow.☆13Jan 1, 2024Updated 2 years ago
- Synthetic data generation for evaluating LLM symbolic and logic reasoning☆23Aug 13, 2026Updated last month
- Lightweight Mini Utility CSS Toolkit☆15Feb 3, 2026Updated 8 months ago
- Folder-structure and some examples for setting up a new ansible-project.☆10Dec 31, 2022Updated 3 years ago
- A step by step implementation of building an AI agent that plays 3d shooting game☆24Jul 16, 2025Updated last year
- A Beginner's Guide to Monetizing Your Python AI Chatbot☆19Apr 22, 2025Updated last year
- Shrink the web for your agents! Token-efficient, fast and fully local web search and research.☆230Updated this week