☆40Apr 5, 2026Updated 6 months ago
Alternatives and similar repositories for qwen3.5-gemma4-moe-flash-mlx-turbo-quant
Users that are interested in qwen3.5-gemma4-moe-flash-mlx-turbo-quant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A memory-first AI agent that remembers why decisions were made — not just the last message. Runs local (Ollama), cloud (Claude · OpenAI ·…☆64Updated this week
- Rust implementation of TurboQuant, PolarQuant, and QJL — zero-overhead vector quantization for semantic search and KV cache compression (…☆26Sep 30, 2026Updated last week
- Flash-MoE sidecar slot-bank runtime for large GGUF MoE models on Apple Silicon — llama.cpp fork☆134Sep 4, 2026Updated last month
- Open-source options gamma exposure (GEX) & positioning dashboard — dealer GEX, max pain, open interest, IV surface. Self-hosted, Docker, …☆79Sep 19, 2026Updated 2 weeks ago
- Lossless DFlash speculative decoding for MLX on Apple Silicon☆787Aug 20, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- 🤵 AIfred-Intelligence — self-hosted multi-agent assistant: debate modes (Symposion/Tribunal), voice (STT + streaming TTS, voice cloning)…☆38Updated this week
- MLX benchmark: Gemma 4 + Qwen 3.5 on Apple Silicon with TurboQuant KV cache☆22Apr 6, 2026Updated 6 months ago
- 🛡️ Programmable Guardrails for LLM Applications in Java. A framework-agnostic toolkit for input/output validation, PII masking, and jail…☆16Apr 15, 2026Updated 5 months ago
- Implementation of websocket server with mio and parser combinators☆16Mar 29, 2022Updated 4 years ago
- ⚡️ The fastest way to run local LLMs on Apple Silicon — sub-second model loads, beats Ollama on throughput, tail latency, and full-respon…☆20Updated this week
- AppleContainerGUI is a native macOS SwiftUI front-end for the container CLI (Apple Container). It provides an organized, visual way to ma…☆28Apr 5, 2026Updated 6 months ago
- Based on the implementation of Google's TurboQuant (ICLR 2026) — Quansloth brings elite KV cache compression to local LLM inference. Qua…☆155May 13, 2026Updated 4 months ago
- Tree-based speculative decoding for Apple Silicon (MLX). ~10-15% faster than DFlash on code, ~1.5x over autoregressive. First MLX port wi…☆145Apr 15, 2026Updated 5 months ago
- Browser for Playdate☆17Dec 15, 2025Updated 9 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- jetworker easy way for community with Web Worker☆10Mar 4, 2023Updated 3 years ago
- Multi-person podcast audio to videocast☆10Sep 28, 2024Updated 2 years ago
- Fastest LLM inference runtime for Apple Silicon☆229Updated this week
- Genome analysis toolkit☆13Apr 23, 2025Updated last year
- my resume☆13Feb 10, 2017Updated 9 years ago
- A hands-on Swift and Metal course for building LLM inference from first principles on Apple silicon, with 48 guided lessons, runnable exe…☆193Jul 17, 2026Updated 2 months ago
- Links to all the source code and solutions I reference in my O'Reilly Introduction to Docker video tutorial☆11Dec 10, 2014Updated 11 years ago
- A HTTP microservice to generate TTS☆16Sep 20, 2026Updated 2 weeks ago
- AI agent toolkit for Compose multiplatform☆15Aug 24, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- My PowerShell Profile Scripts and Modules☆13Jun 19, 2026Updated 3 months ago
- Spy on our Slack room☆12Feb 25, 2016Updated 10 years ago
- ☆37Aug 19, 2026Updated last month
- Docs site for tuya-panel-kit☆13Sep 20, 2026Updated 2 weeks ago
- LudoGL - Magic Box☆16Jun 27, 2025Updated last year
- htop for AI coding agents — monitor token usage, costs, and workflows across Claude Code, Cursor, Kiro, Codex, and Copilot☆49Apr 17, 2026Updated 5 months ago
- A modern CLI for Tenable.io written in Go☆14Nov 28, 2020Updated 5 years ago
- Bigint ID or primary key generator inspired by Twitter's Snowflake and Sonyflake.☆15Oct 29, 2021Updated 4 years ago
- Local AI runtime for training & running small LLMs directly on Apple Neural Engine (ANE). No CoreML. No Metal. Offline, on-device fine-tu…☆129Aug 23, 2026Updated last month
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Elixir-based Event Source server-side implementation using Phoenix Pubsub☆18Nov 25, 2020Updated 5 years ago
- AI agent runtime for Rust — type-safe state, multi-protocol serving, plugin extensibility.☆98Sep 4, 2026Updated last month
- Control your VLC with your IRC☆11Apr 1, 2017Updated 9 years ago
- Web-based e-library platform with bookshelf UI, EPUB reader, and social features☆17Sep 29, 2026Updated last week
- Open-source engine for Skyrim, built with Bevy in Rust☆331Updated this week
- Minecraft Mod. Help your friends back up after they die (if you can make it in time)☆14Sep 16, 2026Updated 3 weeks ago
- ☆15Sep 19, 2024Updated 2 years ago