3.34× faster inference on Apple Silicon — native MLX port of DFlash speculative decoding
☆19Apr 11, 2026Updated 5 months ago
Alternatives and similar repositories for mlx-dflash
Users that are interested in mlx-dflash are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆18Dec 1, 2025Updated 9 months ago
- Lightweight API Specification for Intelligent Systems☆16Sep 8, 2026Updated last week
- a tool deployed on AWS Lambda to help checking internship slot (HCMUT), Public version☆10May 26, 2025Updated last year
- DocFinder is a local-first indexing and searching documents using semantic embeddings stored in SQLite. Everything runs on your machine, …☆32Updated this week
- A Chrome extension that enables virtual fashion try-on and model swap using FASHN AI. Hover over fashion images on any website to: (1) tr…☆22Aug 14, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- A robust Python toolkit for converting video/audio content into accurate, multilingual subtitles using WhisperX for transcription and Goo…☆29Updated this week
- A bytebot variant that uses Holo 1.5 7b to control the desktop☆25Nov 4, 2025Updated 10 months ago
- ☆42Jul 13, 2026Updated 2 months ago
- A MCP stdio toolpack for local LLMs☆34Apr 6, 2026Updated 5 months ago
- An educational Rust project for exporting and running inference on Qwen3 LLM family☆44Aug 3, 2025Updated last year
- A miniaturized version of the Kimi-K2 model optimized for deployment on single H100 GPUs.☆36Jul 16, 2025Updated last year
- Nats jetstream queue adapter with client for Laravel framework☆36Jun 29, 2026Updated 2 months ago
- open-source, local-first AI CLI for developers. It connects to your local Ollama models or remote providers like OpenAI, Anthropic, Gemin…☆61Jul 2, 2026Updated 2 months ago
- Some custom mode-setting system prompts for Roo Code☆20Mar 3, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Fulloch - The Fully Local Home Voice Assistant☆153Updated this week
- Collection of my Google DevFest/Codelab in 2025☆18Apr 11, 2026Updated 5 months ago
- Speculative Decoding Implementations: MTP, EAGLE-3, Medusa-1, PARD, Draft Models, N-gram and Suffix Decoding from scratch☆15May 2, 2026Updated 4 months ago
- Zero-instrumentation LLM API and MCP tracer for your agents powered by eBPF — latency, tokens, and tool use in realtime☆18Mar 16, 2026Updated 6 months ago
- [KDD'25] Code of "Bridging Textual-Collaborative Gap through Semantic Codes for Sequential Recommendation".☆18Jul 3, 2026Updated 2 months ago
- npm package template with typescript and tsup☆11Nov 27, 2025Updated 9 months ago
- Production-grade agent orchestration for Claude Code - 11 agents, 46 MCP tools, SQLite+FTS5, drift detection, consensus checkpoints☆53Jun 8, 2026Updated 3 months ago
- Outlier Detection with AI + ML☆16Sep 12, 2025Updated last year
- Train Llama 3 models from scratch. Any scale, any personality. By Arianna Method.☆52May 4, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆54Feb 19, 2026Updated 6 months ago
- Tiny, composable Atomic CSS engine☆13Apr 1, 2022Updated 4 years ago
- ☆11Feb 28, 2022Updated 4 years ago
- Load and run Llama from safetensors files in C☆16Oct 24, 2024Updated last year
- Vite utility for vue3 server side rendering☆10Jul 21, 2026Updated last month
- Composition of Multimodal Language Models From Scratch☆15Aug 16, 2024Updated 2 years ago
- A hands-on guide for AI builders: make your own RTX 4090D/5090 GPU server that’s fast and efficient.☆17Aug 19, 2025Updated last year
- Local modular AI assistant with speech, vision, and robotics support. Uses Qwen3-VL-4B in LM Studio.☆54Jan 9, 2026Updated 8 months ago
- Implement rest api service for manipulating blog contents using FastAPI in Python☆12Feb 14, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Agentic BYOK Browser-Based Website Builder☆60Updated this week
- Desktop app with Compose Multiplatform to use Qwen3-TTS with an UI.☆76Aug 2, 2026Updated last month
- Code and Data for Evaluating the Evaluators☆16Aug 20, 2025Updated last year
- Is strawberry a fruit or a vegetable?☆54Jun 10, 2026Updated 3 months ago
- ☆16Feb 1, 2025Updated last year
- General Tool-calling API Proxy☆60Mar 26, 2026Updated 5 months ago
- Script Execution service☆13Nov 21, 2016Updated 9 years ago