Flash weight streaming for MLX: run massive models larger than your RAM on Apple Silicon.
☆122Jun 13, 2026Updated last month
Alternatives and similar repositories for mlx-flash
Users that are interested in mlx-flash are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆36Mar 30, 2026Updated 3 months ago
- ☆218Mar 24, 2026Updated 4 months ago
- Running a big model on a small laptop☆3,994Mar 19, 2026Updated 4 months ago
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆727May 19, 2026Updated 2 months ago
- ANE (Apple Neural Engine) CostModel profiler for CoreML models☆36Apr 9, 2026Updated 3 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 4 months ago
- PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal ac…☆306Jun 5, 2026Updated last month
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆117Jul 15, 2026Updated 2 weeks ago
- Implementation of E2-TTS, "Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS", in MLX☆21Oct 8, 2024Updated last year
- Running a big model on a small laptop☆55Mar 28, 2026Updated 4 months ago
- Provides a frame iterator for videos by using ffmpeg. Decodes images using the image crate.☆12Mar 31, 2021Updated 5 years ago
- ☆86Mar 3, 2026Updated 4 months ago
- MLX Model Quantization Toolkit - Comprehensive collection of Jupyter notebooks for converting and quantizing large language models usi…☆16Aug 16, 2025Updated 11 months ago
- ☆21Oct 9, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Shared personal notes created while working with the Apple MLX machine learning framework☆24Dec 12, 2025Updated 7 months ago
- MLX version of DINO DETR☆16Dec 26, 2024Updated last year
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆757Jun 11, 2026Updated last month
- A repo of useful MLX skills.☆87Jan 25, 2026Updated 6 months ago
- An MLX port of Meta's Coconut reasoning model☆16Sep 2, 2025Updated 10 months ago
- [WIP] Open source implementation of Apple's UXKit PrivateFramework☆27Jun 15, 2026Updated last month
- A wrapper around Apple's Container framework that looks and feels like docker compose☆17Jun 19, 2025Updated last year
- A Custom Calendar Designed in SwiftUI☆13Feb 1, 2022Updated 4 years ago
- Este repositorio sirve como guia para la tecnologia vista en el curso. El codigo presente no es absoluto y se puede considerar como plagi…☆15Mar 16, 2026Updated 4 months ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Run Time Series Foundation Models on Apple Silicon☆35Feb 27, 2026Updated 5 months ago
- ☆39Jun 5, 2026Updated last month
- Framework for neural network model creating, training and using in predicting written with use of Metal Performance Shaders☆31Nov 5, 2024Updated last year
- gmacFTP — Free, open-source, native dual-pane FTP/FTPS/SFTP client for macOS. Universal for Apple Silicon and Intel, built in Rust with S…☆20Updated this week
- vibevoice real time 0.5B swift port☆31Dec 12, 2025Updated 7 months ago
- ☆68Jun 4, 2026Updated last month
- Bottom sheet popover built with Swift & UIKit☆12Oct 12, 2020Updated 5 years ago
- macOS audio loopback driver☆16Oct 16, 2020Updated 5 years ago
- The ultimate training toolkit for finetuning diffusion models☆34Jan 22, 2026Updated 6 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Local AI runtime for training & running small LLMs directly on Apple Neural Engine (ANE). No CoreML. No Metal. Offline, on-device fine-tu…☆110Mar 6, 2026Updated 4 months ago
- MLX Model Manager unifies loading and inferencing with LLMs and VLMs.☆102Jan 30, 2025Updated last year
- Train and run transformers directly on Apple's Neural Engine in Swift bypass coreml entirely☆156Jul 16, 2026Updated last week
- Train Large Language Models on MLX.☆402Jul 21, 2026Updated last week
- Chatons is a desktop AI workspace for coding and project workflows: it lets you chat with multiple AI providers, pick scoped or full mod…☆18May 27, 2026Updated 2 months ago
- A lightweight CLI and local API server to create, run and manage macOS and Linux virtual machines (VMs) natively on Apple Silicon.☆15Feb 2, 2025Updated last year
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆146Apr 15, 2026Updated 3 months ago