Official Implementation For PolarQuant
☆46Apr 2, 2026Updated 4 months ago
Alternatives and similar repositories for PolarQuant
Users that are interested in PolarQuant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for Paper 'DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards'☆17May 21, 2026Updated 2 months ago
- The official code implementation of the ACL2025 paper “A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with…☆18Jul 12, 2025Updated last year
- Official implementation for "Mixture of In-Context Experts Enhance LLMs’ Awareness of Long Contexts" (Accepted by Neurips2024)☆14Jan 7, 2025Updated last year
- [ICLR 2026] Official PyTorch implementation for "ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding"☆63Dec 26, 2025Updated 7 months ago
- [EMNLP 2025🔥] UNComp: Can Matrix Entropy Uncover Sparsity? -- A Compressor Design from an Uncertainty-Aware Perspective☆20Jan 7, 2026Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The code for paper Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models.☆13Apr 10, 2024Updated 2 years ago
- Issue tracker for https://alphaxiv.org☆24Oct 13, 2025Updated 9 months ago
- EOQ: Entropy-Optimal Quantization for LLMs. 11-41% smaller than GGUF Q4_K_M with near-FP16 perplexity.☆46Mar 31, 2026Updated 4 months ago
- Algorithm-System Co-design: accurate and efficient 2-bit KV cache quantization for LLM Inference.☆19May 20, 2026Updated 2 months ago
- PyTorch code for full quantization of DNN using BCGD☆14Jul 24, 2019Updated 7 years ago
- Code and models for EMNLP 2024 paper "WPO: Enhancing RLHF with Weighted Preference Optimization"☆41Sep 24, 2024Updated last year
- Code for paper: Long cOntext aliGnment via efficient preference Optimization☆26Oct 10, 2025Updated 10 months ago
- ☆13Mar 5, 2025Updated last year
- An Tensorflow.keras implementation of Same, Same But Different - Recovering Neural Network Quantization Error Through Weight Factorizatio…☆10Dec 18, 2019Updated 6 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A PyTorch Implementation of Feature Boosting and Suppression☆18Sep 14, 2020Updated 5 years ago
- [ICLR 2026] SERE: Similarity-Based Expert Re-routing for Efficient Batch Decoding in MoE Models☆19Feb 4, 2026Updated 6 months ago
- Code for the AAAI 2024 Oral paper "OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Model…☆72Mar 7, 2024Updated 2 years ago
- The official code implementation for paper "PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs"☆29May 24, 2025Updated last year
- Code needed to reproduce results from my ICLR 2019 paper on fixed-point quantization of the backprop algorithm.☆10Jan 24, 2019Updated 7 years ago
- ☆11Dec 8, 2022Updated 3 years ago
- ThinK: Thinner Key Cache by Query-Driven Pruning☆30Jun 2, 2026Updated 2 months ago
- HTTP proxy with service discovery capabilities based on consul☆17Aug 3, 2017Updated 9 years ago
- An experimentation platform for LLM inference optimisation☆36Sep 19, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- High-Ratio Vector Quantization☆15Feb 3, 2026Updated 6 months ago
- [TMLR] Official PyTorch implementation of paper "Efficient Quantization-aware Training with Adaptive Coreset Selection"☆39Aug 20, 2024Updated last year
- A macOS .mobileconfig generator for installing fonts on iOS device. 一个 macOS 上的 .mobileconfig 配置生成器,用于给 iOS 设备安装字体。☆12Jun 24, 2020Updated 6 years ago
- [ICML 2025] Retraining-Free Merging of Sparse MoE via Hierarchical Clustering☆25Oct 26, 2025Updated 9 months ago
- Web page for exponent pair database☆14Aug 5, 2024Updated 2 years ago
- Unit Scaling demo and experimentation code☆16Mar 12, 2024Updated 2 years ago
- WoW: A Window-to-Window Incremental Index for Range-Filtering Approximate Nearest Neighbor Search, SIGMOD 2026☆17Sep 25, 2025Updated 10 months ago
- ☆22Sep 29, 2025Updated 10 months ago
- [ICML 2021] "Double-Win Quant: Aggressively Winning Robustness of Quantized DeepNeural Networks via Random Precision Training and Inferen…☆16Feb 13, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This Python bot automates trading on Polymarket by comparing interest rate probabilities from the bond market and Polymarket. It detects …☆12Jan 19, 2025Updated last year
- This repository provides code source used in the paper: A Mean Field Theory of Quantized Deep Networks: The Quantization-Depth Trade-Off☆13May 30, 2019Updated 7 years ago
- Code for paper 'Minimizing FLOPs to Learn Efficient Sparse Representations' published at ICLR 2020☆20Feb 14, 2020Updated 6 years ago
- ☆16Jun 17, 2026Updated last month
- ☆17Jun 27, 2026Updated last month
- [ICML 2026]A framework to compare low-bit integer and float-point formats☆82May 6, 2026Updated 3 months ago
- BMXNet: An Open-Source Binary Neural Network Implementation Based on MXNet☆17Dec 5, 2018Updated 7 years ago