REAP expert pruning for MoE LLMs on Apple Silicon via MLX
☆58Mar 16, 2026Updated 4 months ago
Alternatives and similar repositories for reap-mlx
Users that are interested in reap-mlx are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆37Jun 12, 2026Updated last month
- Minimalist repo to do mlx vlm constrained decoding in batch mode☆18Apr 11, 2026Updated 3 months ago
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆62Apr 18, 2026Updated 3 months ago
- Train Embedding Models on MLX.☆17Jun 2, 2026Updated 2 months ago
- Moshi-Finetune-MLX lets you fine-tune Moshi (Native, Real-Time, Speech-to-Speech) models all on Apple Silicon.☆26Apr 21, 2026Updated 3 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Triton‑style kernel toolkit for MLX plus a small upstream incubator: prototype, benchmark, and upstream fusions for Apple Silicon☆48Mar 31, 2026Updated 4 months ago
- An MLX port of Meta's Coconut reasoning model☆16Sep 2, 2025Updated 11 months ago
- ☆36Mar 30, 2026Updated 4 months ago
- Find the hidden meaning of LLMs☆42Nov 13, 2025Updated 8 months ago
- Minimal Claude Code alternative powered by MLX☆47Jan 11, 2026Updated 6 months ago
- ☆221Mar 24, 2026Updated 4 months ago
- ☆18May 27, 2025Updated last year
- ollama like cli tool for MLX models on huggingface (pull, rm, list, show, serve etc.)☆151Updated this week
- this repo has all official MLX-LM-LoRA example notebooks for training on Apple Silicon☆36Apr 23, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- How much experts do we need to serve a model?☆152Mar 18, 2026Updated 4 months ago
- Mini AI Developer☆20Mar 17, 2026Updated 4 months ago
- ☆179Mar 30, 2026Updated 4 months ago
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆118Jul 15, 2026Updated 3 weeks ago
- MLX Implementation of Recursive Reasoning with Tiny Networks☆79Oct 11, 2025Updated 9 months ago
- A native Mac App for LLM fine-tuning on Apple Silicon — fully on-device, fully open source.☆254Jul 16, 2026Updated 3 weeks ago
- Exploratory stuff on RLMs for video workflows☆35Apr 20, 2026Updated 3 months ago
- MLX-Video is the best package for inference and finetuning of Image-Video-Audio generation models on your Mac using MLX.☆284May 13, 2026Updated 2 months ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆758Jun 11, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- The ultimate training toolkit for finetuning diffusion models☆34Jan 22, 2026Updated 6 months ago
- This repo maintains a 'cheat sheet' for LLMs that are undertrained on mlx☆33Mar 12, 2026Updated 4 months ago
- Running a big model on a small laptop☆56Mar 28, 2026Updated 4 months ago
- ☆50Updated this week
- CLI for Recursive Language Models (arXiv:2512.24601)☆205Jun 17, 2026Updated last month
- Exact speculative decoding on Apple Silicon, powered by MLX.☆382Apr 20, 2026Updated 3 months ago
- MLX binary vectors and associated algorithms.☆14Mar 13, 2025Updated last year
- MLX-Embeddings is the best package for running Vision and Language Embedding models locally on your Mac using MLX.☆425May 13, 2026Updated 2 months ago
- mlx-lm server wrapper for agentic harness☆21Jan 26, 2026Updated 6 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Train Large Language Models on MLX.☆407Jul 21, 2026Updated 2 weeks ago
- On-device semantic search over Apple WWDC 2025 docs using MLX embeddings — SwiftUI app (WWDC OMT 2025)☆76Jun 12, 2025Updated last year
- import documents for LLMs☆48Jul 7, 2026Updated last month
- ☆13Jun 29, 2024Updated 2 years ago
- Home for the Development of MLX Vulkan backend☆38Jul 16, 2026Updated 3 weeks ago
- alternative way to calculating self attention☆18May 25, 2024Updated 2 years ago
- 3x decode TPS increase On Qwen 3.6 27B @ temp 0.6 | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.☆1,144Updated this week