Official Repository for Task-Circuit Quantization
☆28Jun 1, 2025Updated last year
Alternatives and similar repositories for TACQ
Users that are interested in TACQ are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- IntructIR, a novel benchmark specifically designed to evaluate the instruction following ability in information retrieval models. Our foc…☆32Jun 13, 2024Updated 2 years ago
- ☆35May 16, 2025Updated last year
- [ICLR 2026] This is the official PyTorch implementation of "QVGen: Pushing the Limit of Quantized Video Generative Models".☆33Feb 11, 2026Updated 6 months ago
- 🎓Automatically Update circult-eda-mlsys-tinyml Papers Daily using Github Actions (Update Every 8th hours)☆10Updated this week
- ☆17May 19, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- KV cache compression via sparse coding☆18Oct 26, 2025Updated 10 months ago
- Improved the performance of 8-bit PTQ4DM expecially on FID.☆11Aug 30, 2023Updated 3 years ago
- ☆16Feb 18, 2024Updated 2 years ago
- ☆15Jun 26, 2024Updated 2 years ago
- Solving some interesting problems using Python and C++☆14Aug 16, 2020Updated 6 years ago
- ICLR 2026☆30May 13, 2026Updated 3 months ago
- [ACM MM2025]: MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization☆44Aug 13, 2025Updated last year
- [EMNLP 2024] CompAct: Compressing Retrieved Documents Actively for Question Answering☆37Sep 20, 2024Updated last year
- An algorithm for weight-activation quantization (W4A4, W4A8) of LLMs, supporting both static and dynamic quantization☆177Nov 26, 2025Updated 9 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- This repository includes the code to download the curated HuggingFace papers into a single markdown formatted file☆16Jul 26, 2024Updated 2 years ago
- ☆16Jan 14, 2025Updated last year
- Fork of Flame repo for training of some new stuff in development☆20Updated this week
- ☆16Nov 25, 2022Updated 3 years ago
- 3D reconstruction and Plane detection using plane-to-plane homography constraints for uncalibrated image pair under Manhattan World Assum…☆17Dec 2, 2019Updated 6 years ago
- Pytorch implementation of "Oscillation-Reduced MXFP4 Training for Vision Transformers" on DeiT Model Pre-training☆41May 4, 2026Updated 3 months ago
- Information Bottleneck in DNN with PyTorch☆15Jul 6, 2023Updated 3 years ago
- Artifact for Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization☆18May 9, 2025Updated last year
- ☆25Sep 4, 2025Updated 11 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official code for the paper "Examining Post-Training Quantization for Mixture-of-Experts: A Benchmark"☆31Jun 30, 2025Updated last year
- Instruction Following Eval☆18Jan 16, 2025Updated last year
- ☆186Jun 22, 2025Updated last year
- Machine Learning Function Approximation: This code implements the fully-connected Deep Neural Network (DNN) architectures considered in t…☆20Oct 27, 2020Updated 5 years ago
- Demo for Qwen2.5-VL-3B-Instruct on Axera device.☆16Sep 3, 2025Updated 11 months ago
- Image Quality Assessment Paper Reading☆15Sep 11, 2022Updated 3 years ago
- DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting☆19Mar 4, 2025Updated last year
- ☆20Mar 6, 2022Updated 4 years ago
- An unofficial implementation of the Infini-gram model proposed by Liu et al. (2024)☆33Jun 19, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICML 2023] SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models☆23Mar 15, 2024Updated 2 years ago
- ☆21Aug 19, 2025Updated last year
- 用于解锁密码加密的pptx文件,使其可以编辑。☆15May 6, 2021Updated 5 years ago
- (AAAI 2026) First-Order Error Matters: Accurate Compensation for Quantized Large Language Models☆17Apr 16, 2026Updated 4 months ago
- Pulp virtual platform☆24Jul 16, 2025Updated last year
- Official Implementation of Paper FOLDER (ICCV2025) and Turbo (ECCV2024)☆15Jun 27, 2025Updated last year
- ☆11Jun 24, 2021Updated 5 years ago