TurboQuant KV cache compression for MLX with fused Metal kernels. 4.6x compression at 98% FP16 speed.
☆114Apr 30, 2026Updated 4 months ago
Alternatives and similar repositories for turboquant-mlx
Users that are interested in turboquant-mlx are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A tiny server to run local inference on MLX model in the style of OpenAI☆13Jan 31, 2024Updated 2 years ago
- GUF to MLX Converter; LM studio advanced tips; Simple text generation using Apple's MLX framework;Enhanced MLX implementation with transf…☆39Updated this week
- ☆21Apr 30, 2026Updated 4 months ago
- Deploy your GGML models to HuggingFace Spaces with Docker and gradio☆37Jun 6, 2023Updated 3 years ago
- Local-first AI agent framework with GUI, memory, web search, personality constructs, speech i/o, tools, skills, CLI & Telegram features —…☆26Mar 20, 2026Updated 6 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon☆275Jun 3, 2026Updated 3 months ago
- ☆15May 17, 2024Updated 2 years ago
- An on device Agent Runtime☆23May 13, 2026Updated 4 months ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆785Aug 20, 2026Updated last month
- An open-source skill for running parallel implementations, reviewing them independently, and selecting or synthesizing the best result.☆32Apr 1, 2026Updated 5 months ago
- A customizable Claude Code setup for knowledge workers: agents, safety hooks, skills, memory, and an Obsidian knowledge base. Clone, run …☆32Jun 23, 2026Updated 2 months ago
- ☆31Sep 1, 2023Updated 3 years ago
- A Model Context Protocol server starter template☆34Feb 11, 2025Updated last year
- Llm wiki in open knowledge format☆20Jul 25, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Universal context usage analyzer for Claude Code - works from any directory in any project☆19Nov 16, 2025Updated 10 months ago
- JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon☆227Sep 4, 2026Updated 2 weeks ago
- ⚡️ The fastest way to run local LLMs on Apple Silicon — sub-second model loads, beats Ollama on throughput, tail latency, and full-respon…☆21Updated this week
- NewsAgent is an enterprise-grade news aggregation agent designed to fetch, query, and summarize news from multiple sources at scale.☆30Oct 13, 2025Updated 11 months ago
- ☆12Aug 1, 2025Updated last year
- REAM: Merging Improves Pruning of Experts in LLMs☆26Apr 16, 2026Updated 5 months ago
- Agentic Runtime Extensible Server. A composable AI agent runtime in Rust based on the Cordis framework☆15Updated this week
- A memory-first AI agent that remembers why decisions were made — not just the last message. Runs local (Ollama), cloud (Claude · OpenAI ·…☆60Updated this week
- An Opencode plugin for managing git worktrees.☆77Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆199Mar 29, 2026Updated 5 months ago
- Accepts a Hugging Face model URL, automatically downloads and quantizes it using Bits and Bytes.☆38Mar 12, 2024Updated 2 years ago
- ☆16May 8, 2025Updated last year
- Experimental Project: What if Gemma had an SDK for web developers to builds games around it as the game brain?☆43May 11, 2026Updated 4 months ago
- Experimental framework taking inspiration from biological systems, combining compression-based architectures, group theory, and symmetry …☆14Nov 13, 2025Updated 10 months ago
- Interface boards for the Neato LDS - the Lidar sensor of the XV-11.☆14Jun 21, 2014Updated 12 years ago
- 1.58 Bit LLM on Apple Silicon using MLX☆304May 10, 2024Updated 2 years ago
- Remote MCP to run OpenCode on Sandboxes☆24Jan 9, 2026Updated 8 months ago
- ☆12Jan 2, 2020Updated 6 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 📦 Unofficial library, CLI tool, and MCP server for the Finnish Tori.fi marketplace.☆20Jun 22, 2026Updated 2 months ago
- An AI bot which helps in day-to-day work☆18Apr 7, 2026Updated 5 months ago
- A local MCP server that gives AI agents fast symbol search, call-graph traversal, and blast-radius analysis over your codebase.☆17Mar 7, 2026Updated 6 months ago
- ☆29Updated this week
- Structured output benchmarks comparing DSPy and BAML with different LLMs☆29Dec 23, 2025Updated 8 months ago
- Shows how to make a macOS app that opens and edits CSV files. It is a SwiftUI project using TableView, fileExporter, fileImporter, and R…☆46Nov 12, 2024Updated last year
- MindBridge is an AI orchestration MCP server that lets any app talk to any LLM — OpenAI, Anthropic, DeepSeek, Ollama, and more — through …☆37Mar 13, 2026Updated 6 months ago