☆37Mar 30, 2026Updated 5 months ago
Alternatives and similar repositories for anemll-flash-mlx
Users that are interested in anemll-flash-mlx are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆132Sep 4, 2026Updated last week
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆61Apr 18, 2026Updated 4 months ago
- Find the hidden meaning of LLMs☆42Nov 13, 2025Updated 10 months ago
- ANE (Apple Neural Engine) CostModel profiler for CoreML models☆36Apr 9, 2026Updated 5 months ago
- DeepSeek V4 Flash specific inference engine. SSD MoE expert paging (slot-bank) + disk KV cache for long agent sessions. Metal-first, narr…☆31Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Multi-LoRA inference server for Apple Silicon -- one base model, many adapters, zero reload☆20Apr 13, 2026Updated 5 months ago
- Flash weight streaming for MLX: run massive models larger than your RAM on Apple Silicon.☆132Jun 13, 2026Updated 3 months ago
- ☆222Mar 24, 2026Updated 5 months ago
- Moshi-Finetune-MLX lets you fine-tune Moshi (Native, Real-Time, Speech-to-Speech) models all on Apple Silicon.☆26Apr 21, 2026Updated 4 months ago
- Local agent infrastructure in one stdlib-only Go binary☆17Updated this week
- REAP expert pruning for MoE LLMs on Apple Silicon via MLX☆58Mar 16, 2026Updated 5 months ago
- MLX Model Quantization Toolkit - Comprehensive collection of Jupyter notebooks for converting and quantizing large language models usi…☆16Aug 16, 2025Updated last year
- Minimal Claude Code alternative powered by MLX☆47Jan 11, 2026Updated 8 months ago
- ☆45Mar 5, 2026Updated 6 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- NetHack as-a-library. Semantic world exploration engine.☆49Updated this week
- A repo of useful MLX skills.☆89Jan 25, 2026Updated 7 months ago
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆766Sep 5, 2026Updated last week
- Running a big model on a small laptop☆57Mar 28, 2026Updated 5 months ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆779Aug 20, 2026Updated 3 weeks ago
- vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon☆275Jun 3, 2026Updated 3 months ago
- Triton‑style kernel toolkit for MLX plus a small upstream incubator: prototype, benchmark, and upstream fusions for Apple Silicon☆53Mar 31, 2026Updated 5 months ago
- ragequit everywhere☆20Nov 14, 2022Updated 3 years ago
- Exact speculative decoding on Apple Silicon, powered by MLX.☆388Apr 20, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 60-95% fewer tokens, same answers☆30May 19, 2026Updated 3 months ago
- Train Embedding Models on MLX.☆17Jun 2, 2026Updated 3 months ago
- Pure MLX implementations of UMAP, t-SNE, PaCMAP, TriMap, DREAMS, CNE, MMAE, and NNDescent for Apple Silicon. Metal GPU for computation an…☆93Mar 20, 2026Updated 5 months ago
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆145Apr 15, 2026Updated 4 months ago
- Chrome Web Store☆19Jul 11, 2024Updated 2 years ago
- ModernBERT model optimized for Apple Neural Engine.☆39Jan 10, 2025Updated last year
- Implementation of Contrastive Neuron Attribution for behavioral detection and steering.☆37Jun 1, 2026Updated 3 months ago
- Test LLMs on real tasks. Compare models side-by-side.☆422Aug 10, 2026Updated last month
- A command-line utility to manage MLX models between your Hugging Face cache and LM Studio.☆89Nov 11, 2025Updated 10 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Libegpu is a library for enumerating eGPU devices & enclosures.☆17Nov 24, 2025Updated 9 months ago
- LLM training on Apple's Neural Engine — native Obj-C, private APIs, zero GPU. Dynamic weight pipeline for training without kernel recompi…☆58Mar 17, 2026Updated 5 months ago
- Utilities to evaluate MLX quantizations☆26Jun 9, 2026Updated 3 months ago
- Run transformers directly on Apple's Neural Engine in Swift bypass coreml entirely☆168Updated this week
- Thoughtful Lightning AI Assistant - Dual-engine system with DeepSeek reasoning and Groq inference, featuring Gradio UI, secure API manage…☆20Jan 22, 2025Updated last year
- ☆12Sep 25, 2022Updated 3 years ago
- Solidity library to query contract ABI in Foundry tests and scripts☆14Apr 26, 2024Updated 2 years ago