☆36Mar 30, 2026Updated 4 months ago
Alternatives and similar repositories for anemll-flash-mlx
Users that are interested in anemll-flash-mlx are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆129Aug 15, 2026Updated last week
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆61Apr 18, 2026Updated 4 months ago
- Find the hidden meaning of LLMs☆42Nov 13, 2025Updated 9 months ago
- ANE (Apple Neural Engine) CostModel profiler for CoreML models☆36Apr 9, 2026Updated 4 months ago
- DeepSeek V4 Flash specific inference engine. SSD MoE expert paging (slot-bank) + disk KV cache for long agent sessions. Metal-first, narr…☆31Jul 18, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Multi-LoRA inference server for Apple Silicon -- one base model, many adapters, zero reload☆20Apr 13, 2026Updated 4 months ago
- Minimalist repo to do mlx vlm constrained decoding in batch mode☆18Apr 11, 2026Updated 4 months ago
- An MLX port of Meta's Coconut reasoning model☆16Sep 2, 2025Updated 11 months ago
- Flash weight streaming for MLX: run massive models larger than your RAM on Apple Silicon.☆126Jun 13, 2026Updated 2 months ago
- ☆223Mar 24, 2026Updated 4 months ago
- Moshi-Finetune-MLX lets you fine-tune Moshi (Native, Real-Time, Speech-to-Speech) models all on Apple Silicon.☆26Apr 21, 2026Updated 4 months ago
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 5 months ago
- MLX Model Quantization Toolkit - Comprehensive collection of Jupyter notebooks for converting and quantizing large language models usi…☆16Aug 16, 2025Updated last year
- Minimal Claude Code alternative powered by MLX☆47Jan 11, 2026Updated 7 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆44Mar 5, 2026Updated 5 months ago
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆744Aug 16, 2026Updated last week
- A repo of useful MLX skills.☆88Jan 25, 2026Updated 6 months ago
- Running a big model on a small laptop☆56Mar 28, 2026Updated 4 months ago
- mlx-lm server wrapper for agentic harness☆21Jan 26, 2026Updated 6 months ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆769Updated this week
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆38Jun 12, 2026Updated 2 months ago
- Convert StableHLO models into Apple Core ML format☆22Updated this week
- vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon☆274Jun 3, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Exact speculative decoding on Apple Silicon, powered by MLX.☆387Apr 20, 2026Updated 4 months ago
- Agentic authentication including SSO examples for all major frameworks☆16Jul 15, 2026Updated last month
- Train Embedding Models on MLX.☆17Jun 2, 2026Updated 2 months ago
- Pure MLX implementations of UMAP, t-SNE, PaCMAP, TriMap, DREAMS, CNE, MMAE, and NNDescent for Apple Silicon. Metal GPU for computation an…☆92Mar 20, 2026Updated 5 months ago
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆146Apr 15, 2026Updated 4 months ago
- Implementation of Contrastive Neuron Attribution for behavioral detection and steering.☆37Jun 1, 2026Updated 2 months ago
- A pure MLX-based training pipeline for fine-tuning LLMs using GRPO on Apple Silicon.☆242Oct 28, 2025Updated 9 months ago
- Test LLMs on real tasks. Compare models side-by-side.☆413Aug 10, 2026Updated last week
- A command-line utility to manage MLX models between your Hugging Face cache and LM Studio.☆88Nov 11, 2025Updated 9 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Libegpu is a library for enumerating eGPU devices & enclosures.☆17Nov 24, 2025Updated 9 months ago
- LLM training on Apple's Neural Engine — native Obj-C, private APIs, zero GPU. Dynamic weight pipeline for training without kernel recompi…☆58Mar 17, 2026Updated 5 months ago
- Utilities to evaluate MLX quantizations☆27Jun 9, 2026Updated 2 months ago
- Apple Neural Engine (ANE) LLM inference engine — reverse-engineered private APIs, Metal GPU shaders, hybrid ANE+GPU+CPU on Apple Silicon.…☆22Mar 5, 2026Updated 5 months ago
- Thoughtful Lightning AI Assistant - Dual-engine system with DeepSeek reasoning and Groq inference, featuring Gradio UI, secure API manage…☆20Jan 22, 2025Updated last year
- ☆12Sep 25, 2022Updated 3 years ago
- Solidity library to query contract ABI in Foundry tests and scripts☆14Apr 26, 2024Updated 2 years ago