☆36Mar 30, 2026Updated 4 months ago
Alternatives and similar repositories for anemll-flash-mlx
Users that are interested in anemll-flash-mlx are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆117Jul 15, 2026Updated 2 weeks ago
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆62Apr 18, 2026Updated 3 months ago
- Find the hidden meaning of LLMs☆42Nov 13, 2025Updated 8 months ago
- ANE (Apple Neural Engine) CostModel profiler for CoreML models☆36Apr 9, 2026Updated 3 months ago
- DeepSeek V4 Flash specific inference engine. SSD MoE expert paging (slot-bank) + disk KV cache for long agent sessions. Metal-first, narr…☆29Jul 18, 2026Updated 2 weeks ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Multi-LoRA inference server for Apple Silicon -- one base model, many adapters, zero reload☆20Apr 13, 2026Updated 3 months ago
- Minimalist repo to do mlx vlm constrained decoding in batch mode☆18Apr 11, 2026Updated 3 months ago
- An MLX port of Meta's Coconut reasoning model☆16Sep 2, 2025Updated 11 months ago
- Flash weight streaming for MLX: run massive models larger than your RAM on Apple Silicon.☆123Jun 13, 2026Updated last month
- ☆218Mar 24, 2026Updated 4 months ago
- Moshi-Finetune-MLX lets you fine-tune Moshi (Native, Real-Time, Speech-to-Speech) models all on Apple Silicon.☆26Apr 21, 2026Updated 3 months ago
- Local agent infrastructure in one stdlib-only Go binary☆15Updated this week
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 4 months ago
- MLX Model Quantization Toolkit - Comprehensive collection of Jupyter notebooks for converting and quantizing large language models usi…☆16Aug 16, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆730May 19, 2026Updated 2 months ago
- A repo of useful MLX skills.☆87Jan 25, 2026Updated 6 months ago
- Running a big model on a small laptop☆55Mar 28, 2026Updated 4 months ago
- mlx-lm server wrapper for agentic harness☆21Jan 26, 2026Updated 6 months ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆756Jun 11, 2026Updated last month
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆37Jun 12, 2026Updated last month
- vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon☆274Jun 3, 2026Updated 2 months ago
- ragequit everywhere☆20Nov 14, 2022Updated 3 years ago
- Exact speculative decoding on Apple Silicon, powered by MLX.☆382Apr 20, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers☆30May 19, 2026Updated 2 months ago
- Train Embedding Models on MLX.☆17Jun 2, 2026Updated 2 months ago
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆145Apr 15, 2026Updated 3 months ago
- Chrome Web Store☆19Jul 11, 2024Updated 2 years ago
- ModernBERT model optimized for Apple Neural Engine.☆38Jan 10, 2025Updated last year
- Implementation of Contrastive Neuron Attribution for behavioral detection and steering.☆33Jun 1, 2026Updated 2 months ago
- A pure MLX-based training pipeline for fine-tuning LLMs using GRPO on Apple Silicon.☆242Oct 28, 2025Updated 9 months ago
- Test LLMs on real tasks. Compare models side-by-side.☆393Jun 16, 2026Updated last month
- A command-line utility to manage MLX models between your Hugging Face cache and LM Studio.☆88Nov 11, 2025Updated 8 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- LLM training on Apple's Neural Engine — native Obj-C, private APIs, zero GPU. Dynamic weight pipeline for training without kernel recompi…☆57Mar 17, 2026Updated 4 months ago
- Apple Neural Engine (ANE) LLM inference engine — reverse-engineered private APIs, Metal GPU shaders, hybrid ANE+GPU+CPU on Apple Silicon.…☆22Mar 5, 2026Updated 4 months ago
- Train and run transformers directly on Apple's Neural Engine in Swift bypass coreml entirely☆157Jul 16, 2026Updated 2 weeks ago
- Thoughtful Lightning AI Assistant - Dual-engine system with DeepSeek reasoning and Groq inference, featuring Gradio UI, secure API manage…☆20Jan 22, 2025Updated last year
- Contracts for EthML- a decentralized AI implementation☆24Dec 14, 2021Updated 4 years ago
- EcoFlow Portable Power Station Integration for Home Assistant☆12Apr 4, 2025Updated last year
- A fully trainable state space model (SSM)☆16Mar 18, 2025Updated last year