mi-optimize is a versatile tool designed for the quantization and evaluation of large language models (LLMs). The library's seamless integration of various quantization methods and evaluation techniques empowers users to customize their approaches according to specific requirements and constraints, providing a high level of flexibility.
☆24Nov 28, 2024Updated last year
Alternatives and similar repositories for MI-optimize
Users that are interested in MI-optimize are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [EMNLP 2025] AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models☆18Apr 29, 2026Updated 4 months ago
- Official PyTorch implementation of "Evolving Search Space for Neural Architecture Search"☆12Aug 18, 2021Updated 5 years ago
- Various test models in WNNX format. It can view with `pip install wnetron && wnetron`☆12Jun 22, 2022Updated 4 years ago
- Improved the performance of 8-bit PTQ4DM expecially on FID.☆11Aug 30, 2023Updated 3 years ago
- ☆19Jul 30, 2021Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ECCV24] MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization☆14Nov 27, 2024Updated last year
- ☆14Apr 24, 2024Updated 2 years ago
- CMix-NN: Mixed Low-Precision CNN Library for Memory-Constrained Edge Devices☆52Mar 19, 2020Updated 6 years ago
- C rewrite of a minimal Python JPEG decoder☆12Jan 2, 2019Updated 7 years ago
- ☆41Dec 15, 2022Updated 3 years ago
- Binary translation in Rust☆12Jun 22, 2020Updated 6 years ago
- BNG Image Format Implementation☆12Sep 19, 2020Updated 6 years ago
- ☆12Nov 17, 2023Updated 2 years ago
- ☆16Nov 25, 2022Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Implements Kalman Filter based tracking for multiple circular objects☆11Jan 8, 2023Updated 3 years ago
- ☆20Nov 27, 2025Updated 10 months ago
- ☆54Jul 18, 2024Updated 2 years ago
- 3D reconstruction and Plane detection using plane-to-plane homography constraints for uncalibrated image pair under Manhattan World Assum…☆17Dec 2, 2019Updated 6 years ago
- Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost☆21May 14, 2026Updated 4 months ago
- This is an implementation of YOLO using LSQ network quantization method.☆22Apr 13, 2022Updated 4 years ago
- INT-Q Extension of the CMSIS-NN library for ARM Cortex-M target☆18Jan 10, 2020Updated 6 years ago
- A simple C++17 header-only library for generating SVG plots☆10Mar 17, 2024Updated 2 years ago
- 🖥️ a toy riscv emulator☆14Oct 20, 2021Updated 4 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆20Mar 6, 2022Updated 4 years ago
- RISC-V instruction encoding/decoding☆12Mar 22, 2023Updated 3 years ago
- RISC-V Static Binary Translator☆18Mar 6, 2019Updated 7 years ago
- This is the official implementation of the paper Taming Reversible Halftoning via Predictive Luminance [TVCG 2023]☆11Feb 18, 2026Updated 7 months ago
- ☆27Apr 15, 2026Updated 5 months ago
- (ICLR 2025) BinaryDM: Accurate Weight Binarization for Efficient Diffusion Models☆25Oct 4, 2024Updated last year
- A profiling library for the Sega Dreamcast☆11Apr 21, 2026Updated 5 months ago
- Asynchronous I/O framework for C with coroutine scheduling☆16Jul 6, 2025Updated last year
- Hinton's Forward-Forward Algorithm for Deep Learning☆10Feb 6, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Evaluation Code repository for the paper "ModuLoRA: Finetuning 3-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers". (2023…☆13Dec 5, 2023Updated 2 years ago
- Running massive simulations using RNNs on CPUs for building bots and all kinds of things.☆12Jun 13, 2021Updated 5 years ago
- Official Repo for SparseLLM: Global Pruning of LLMs (NeurIPS 2024)☆71Mar 27, 2025Updated last year
- Simple C library for safely handling utf8 strings☆16Nov 30, 2014Updated 11 years ago
- Methods of Self Calibration☆20Jun 10, 2019Updated 7 years ago
- TensorRT is a C++ library for high performance inference on NVIDIA GPUs and deep learning accelerators.☆21Mar 7, 2024Updated 2 years ago
- ☆33Sep 3, 2025Updated last year