TurboQuant KV cache compression for MLX with fused Metal kernels. 4.6x compression at 98% FP16 speed.
☆112Apr 30, 2026Updated 2 months ago
Alternatives and similar repositories for turboquant-mlx
Users that are interested in turboquant-mlx are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Cloudflare-native AI agent — 13 tools, codemode, 5-layer memory, self-learning, multimodal I/O. Telegram, Discord & WhatsApp bots. Web se…☆25Jul 7, 2026Updated 2 weeks ago
- A tiny server to run local inference on MLX model in the style of OpenAI☆13Jan 31, 2024Updated 2 years ago
- ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache…☆723May 19, 2026Updated 2 months ago
- GUF to MLX Converter; LM studio advanced tips; Simple text generation using Apple's MLX framework;Enhanced MLX implementation with transf…☆32Jul 9, 2026Updated last week
- OpenClaw AI Agent Swarm Dashboard☆33Apr 6, 2026Updated 3 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆21Apr 30, 2026Updated 2 months ago
- Deploy your GGML models to HuggingFace Spaces with Docker and gradio☆37Jun 6, 2023Updated 3 years ago
- Rust implementation of TurboQuant vector quantization (ICLR 2026, Google Research)☆21Mar 26, 2026Updated 3 months ago
- The honest macOS tune-up — diagnoses what's actually slow, fixes only what's safe & reversible, scores your Mac 0-100. One bash script, z…☆20Jun 13, 2026Updated last month
- Rust bindings for Apple's FoundationModels.framework☆22Updated this week
- Local-first AI agent framework with GUI, memory, web search, personality constructs, speech i/o, tools, skills, CLI & Telegram features —…☆23Mar 20, 2026Updated 4 months ago
- Python toolkit for SIM/eSIM and eUICC work: SCP03, SCP80, SCP11 (relay, local, eIM), SAIP profile packages, and a simulated UICC/eUICC en…☆16Jul 3, 2026Updated 2 weeks ago
- Full-stack OpenTelemetry observability for Apache Spark☆16Feb 28, 2026Updated 4 months ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆753Jun 11, 2026Updated last month
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Pure Rust implementation of Google's TurboQuant (ICLR 2026) — KV cache compression for LLMs☆39Apr 19, 2026Updated 3 months ago
- An on device Agent Runtime☆23May 13, 2026Updated 2 months ago
- A customizable Claude Code setup for knowledge workers: agents, safety hooks, skills, memory, and an Obsidian knowledge base. Clone, run …☆28Jun 23, 2026Updated 3 weeks ago
- 🧠 Bio-Agent OS: 🇻🇳 Bio-Inspired Memory Framework for AI Agents (OpenClaw/ERP). Researched & Developed by Dev Tuan Anh Ha (Locaith Solu…☆20Apr 21, 2026Updated 3 months ago
- ☆31Sep 1, 2023Updated 2 years ago
- ⚡️ The fastest way to run local LLMs on Apple Silicon — sub-second model loads, beats Ollama on throughput, tail latency, and full-respon…☆16Jul 13, 2026Updated last week
- Build real software step-by-step with Claude, Codex, OpenCode, Gemini, Cursor, and other agents☆15Mar 29, 2026Updated 3 months ago
- fastest runtime for apple silicon.☆88Apr 16, 2026Updated 3 months ago
- Universal context usage analyzer for Claude Code - works from any directory in any project☆20Nov 16, 2025Updated 8 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon☆213Updated this week
- import your conversation history into AI 🫚☆16Jul 14, 2026Updated last week
- NewsAgent is an enterprise-grade news aggregation agent designed to fetch, query, and summarize news from multiple sources at scale.☆29Oct 13, 2025Updated 9 months ago
- A memory-first AI agent that remembers why decisions were made — not just the last message. Runs local (Ollama), cloud (Claude · OpenAI ·…☆51Updated this week
- ☆40Mar 12, 2026Updated 4 months ago
- Agentic AI server in Rust. Multi-provider LLM routing, tool calling, RAG, MCP, multi-tenant workflows.☆16Jun 18, 2026Updated last month
- Metal Flash Attention for MLX☆19Jul 14, 2026Updated last week
- An Opencode plugin for managing git worktrees.☆75Mar 25, 2026Updated 3 months ago
- MarrowScript compiler. Welcome to deterministic typed LLM orchestration as a compile-time concern☆32May 21, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A monitoring tool to gather infrastructure network information☆24Updated this week
- A self-hosted AI toolkit running locally via Docker Compose, bundling an LLM gateway, workflow automation, and a chat UI — all backed by …☆16May 17, 2026Updated 2 months ago
- High-performance late-interaction retrieval engine for on-prem AI. ColBERT/ColPali multi-vector search with Rust fused MaxSim, Triton GPU…☆17Jul 6, 2026Updated 2 weeks ago
- ☆22Jun 19, 2026Updated last month
- ☆16May 8, 2025Updated last year
- For inferring and serving local LLMs using the MLX framework☆115Mar 24, 2024Updated 2 years ago
- 1.58 Bit LLM on Apple Silicon using MLX☆294May 10, 2024Updated 2 years ago