Turbo1Bit: Combining 1-bit LLM weights (Bonsai) with TurboQuant KV cache compression for maximum inference efficiency. 4.2x KV cache compression + 16x weight compression = ~10x total memory reduction.
☆31Apr 2, 2026Updated 4 months ago
Alternatives and similar repositories for Turbo1bit
Users that are interested in Turbo1bit are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Ultra-Sparse Adaptation of 1-Bit LLMs via XOR Patches☆87Aug 6, 2026Updated 3 weeks ago
- SubRISC: Simple Instruction-Set Computer for IoT edge devices☆16Jun 27, 2018Updated 8 years ago
- Headless terminal emulator CLI powered by libghostty-vt☆18Apr 8, 2026Updated 4 months ago
- The fastest inference framework to run AI on CPUs☆19Jun 12, 2026Updated 2 months ago
- Dice Language Support for VS Code☆10Sep 29, 2020Updated 5 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A simple library to better manage AI prompts in your Rust code.☆14Apr 29, 2024Updated 2 years ago
- ChatGPT over SMS using Twilio Programmable Messaging, OpenAI API, Flask☆13Jan 22, 2023Updated 3 years ago
- Test your local LLMs on the AIME problems☆39Jun 7, 2025Updated last year
- Cogito Studio: All-in-One Workspace AI☆19Jul 1, 2026Updated last month
- Probabilistic Circuits in Julia☆10Dec 27, 2023Updated 2 years ago
- A tiny nearest-neighbor embedding database written in C☆19Feb 18, 2026Updated 6 months ago
- Gemma 4 31B Abliterated — quality-preserving guardrail removal for Google's most capable open model. Apache 2.0. Runs on Apple Silicon vi…