Minimalist repo to do mlx vlm constrained decoding in batch mode
☆18Apr 11, 2026Updated 4 months ago
Alternatives and similar repositories for mlx-vlm-batch-outlines
Users that are interested in mlx-vlm-batch-outlines are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The best benchmark for LLMs on Apple's MLX framework knowledge and coding tasks.☆38Jun 12, 2026Updated 2 months ago
- DFlash block-diffusion speculative decoding running on Apple Silicon via MLX, with an ANE execution path that explores heterogeneous acce…☆61Apr 18, 2026Updated 4 months ago
- ☆36Mar 30, 2026Updated 4 months ago
- Moshi-Finetune-MLX lets you fine-tune Moshi (Native, Real-Time, Speech-to-Speech) models all on Apple Silicon.☆26Apr 21, 2026Updated 4 months ago
- Notebook Terminal UI and Headless Renderer☆29Apr 19, 2026Updated 4 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆146Apr 15, 2026Updated 4 months ago
- Open source, clear, transparent, real world llm benchmarks☆71Updated this week
- A repo of useful MLX skills.☆88Jan 25, 2026Updated 6 months ago
- 🐙 HRM-MLX implements the Hierarchical Reasoning Model for Apple Silicon with 27M params, enabling fast multi-timescale reasoning on 1k s…☆24Updated this week
- Multi-LoRA inference server for Apple Silicon -- one base model, many adapters, zero reload☆20Apr 13, 2026Updated 4 months ago
- Find the hidden meaning of LLMs☆42Nov 13, 2025Updated 9 months ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆769Updated this week
- this repo has all official MLX-LM-LoRA example notebooks for training on Apple Silicon☆38Apr 23, 2026Updated 4 months ago
- Super fast local inferencing for common NLP tasks on technical text☆54Jun 16, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Minimal Claude Code alternative powered by MLX☆47Jan 11, 2026Updated 7 months ago
- A simple library for generating instruction tuning datasets locally☆107Updated this week
- Thoughtful Lightning AI Assistant - Dual-engine system with DeepSeek reasoning and Groq inference, featuring Gradio UI, secure API manage…☆20Jan 22, 2025Updated last year
- Wrangler Compatible Cloudflare Deployment API☆19Mar 27, 2026Updated 4 months ago
- ☆23Jul 16, 2026Updated last month
- BH hackathon☆14Apr 4, 2024Updated 2 years ago
- Adaptive Precision for EXpert Models: MoE-aware mixed-precision quantization☆447Aug 17, 2026Updated last week
- Pipeline parallel training on Apple Silicon!☆33Dec 5, 2025Updated 8 months ago
- A pure MLX-based training pipeline for fine-tuning LLMs using GRPO on Apple Silicon.☆242Oct 28, 2025Updated 9 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Flash-MoE iOS — Run massive MoE models on iPhone☆48Mar 23, 2026Updated 5 months ago
- Rust-native hybrid training & inference engine for Apple Neural Engine + Metal GPU☆179Apr 3, 2026Updated 4 months ago
- Exact speculative decoding on Apple Silicon, powered by MLX.☆387Apr 20, 2026Updated 4 months ago
- Run LLMs with MLX☆15May 1, 2026Updated 3 months ago
- A beautiful desktop application for Local-NotebookLM that provides an intuitive UI for converting PDFs to engaging audio content.☆23Apr 1, 2026Updated 4 months ago
- Grounded reasoning agent: Falcon Perception + Gemma 4 VLM on Apple Silicon☆28Apr 5, 2026Updated 4 months ago
- OpenSource deployment made easy☆10Jun 13, 2015Updated 11 years ago
- An intelligent load balancer for LM Studio that distributes requests across multiple loaded language models, optimizing resource utilizat…☆21Oct 13, 2025Updated 10 months ago
- Links to all the source code and solutions I reference in my O'Reilly Introduction to Docker video tutorial☆11Dec 10, 2014Updated 11 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Examples for using the SiLLM framework for training and running Large Language Models (LLMs) on Apple Silicon☆16May 8, 2025Updated last year
- ☆18Jan 29, 2026Updated 6 months ago
- ☆15Jun 13, 2025Updated last year
- This repo maintains a 'cheat sheet' for LLMs that are undertrained on mlx☆33Mar 12, 2026Updated 5 months ago
- A CLI secrets store & dispatch tool built for AI agents. Authy stores encrypted secrets locally and dispatches them to agents with polic…☆24Feb 26, 2026Updated 5 months ago
- Local LLM Testing & Benchmarking for Apple Silicon☆202Updated this week
- A simple NextJS MCP client with sensible keybindings☆38Feb 12, 2026Updated 6 months ago