Official implementation of LittleBit (NeurIPS 2025) and its follow-up LittleBit-2 (ICML 2026)
☆222Oct 9, 2026Updated this week
Alternatives and similar repositories for LittleBit
Users that are interested in LittleBit are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Codebase for the Progressive Mixed-Precision Decoding paper.☆22Jul 15, 2025Updated last year
- ☆22Jun 25, 2025Updated last year
- [ASPLOS 2026] M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization.☆17Jan 29, 2026Updated 8 months ago
- Efficient LLM Inference Acceleration using Prompting☆50Oct 22, 2024Updated last year
- Official PyTorch implementation of QwT—“Quantization without Tears” (CVPR 2025): fast, accurate, and hassle-free post-training network qu…☆33Sep 30, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official implementation of LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents☆32Sep 24, 2026Updated 2 weeks ago
- The official code implementation for paper "PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs"☆30May 24, 2025Updated last year
- ☆30Jan 20, 2026Updated 8 months ago
- ☆13Feb 17, 2025Updated last year
- EQ-Net [ICCV 2023]☆32Aug 15, 2023Updated 3 years ago
- Port of Facebook's LLaMA model in C/C++☆13Mar 19, 2023Updated 3 years ago
- Piecewise-Affine Regularized Quantization☆20Feb 5, 2026Updated 8 months ago
- Mic-controlled mouse clicks☆17Oct 6, 2025Updated last year
- ☆19Feb 4, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Milk-V Duo. Access to Internet throw USB RNDIS connection to host machine☆16Jan 11, 2024Updated 2 years ago
- Work in progress.☆82Nov 25, 2025Updated 10 months ago
- GUI tool to QLoRA/LoRA-fine-tune LLMs and deploy to Ollama. Broad GPU support (NVIDIA/AMD/Intel/Apple) + CPU fallback.☆15Feb 18, 2026Updated 7 months ago
- Self-hosted voice for coding agents. Talk from any browser or a Telegram call, interrupt mid-sentence, clone any voice, and hand real wor…☆41Updated this week
- Metrics for evaluating biological sequence design☆16Sep 1, 2026Updated last month
- [ICML 2024] Sparse Model Inversion: Efficient Inversion of Vision Transformers with Less Hallucination☆15Apr 29, 2025Updated last year
- Verifying the optimization phases of the GraalVM compiler☆15Sep 28, 2026Updated last week
- Run BitNet b1.58 ternary LLMs with WebGPU — in browsers and native apps☆23Oct 2, 2026Updated last week
- ☆15Jul 25, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆11Apr 5, 2023Updated 3 years ago
- [NeurIPS 2025] Official PyTorch implementation of "Token Bottleneck: One Token to Remember Dynamics"☆32Feb 2, 2026Updated 8 months ago
- ☆16Dec 9, 2023Updated 2 years ago
- [ICCV 2025] Task-Specific Zero-shot Quantization-Aware Training for Object Detection☆31Sep 26, 2025Updated last year
- Source code of our TNNLS paper "Boosting Convolutional Neural Networks with Middle Spectrum Grouped Convolution"☆12Apr 14, 2023Updated 3 years ago
- ☆17Aug 7, 2026Updated 2 months ago
- SMART introduces a novel test-time framework where Small Language Models (SLMs) reason step-by-step, and Large Language Models (LLMs) pro…☆12Jul 9, 2025Updated last year
- ☆18Jun 14, 2025Updated last year
- The official code for "Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation" | [MM2…☆14Dec 7, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆16Aug 19, 2024Updated 2 years ago
- Prepare for DeekSeek R1 inference: Benchmark CPU, DRAM, SSD, iGPU, GPU, ... with efficient code.☆73Feb 2, 2025Updated last year
- ☆12Jul 30, 2025Updated last year
- Python tools for working with LAMMPS files☆16Updated this week
- Fork of Flame repo for training of some new stuff in development☆20Aug 27, 2026Updated last month
- ☆11Sep 20, 2024Updated 2 years ago
- A simple extension that uses Bark Text-to-Speech for audio output☆33Nov 3, 2023Updated 2 years ago