[ICLR2026] The first W4A4KV4 quantized + 50% sparse LLMs!
☆33Jan 26, 2026Updated 6 months ago
Alternatives and similar repositories for OBR
Users that are interested in OBR are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Minute-long video generation at 24FPS.☆69Mar 28, 2026Updated 4 months ago
- [ICCV2025]Generate one 2K image on single 24GB 3090 GPU!☆88Sep 8, 2025Updated 11 months ago
- REAM: Merging Improves Pruning of Experts in LLMs☆23Apr 16, 2026Updated 4 months ago
- [ICLR2025]: OSTQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitt…☆95Apr 8, 2025Updated last year
- ☆118Feb 26, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official implementation for LaCo (EMNLP 2024 Findings)☆22Oct 3, 2024Updated last year
- [ICLR 2026] Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models☆31Mar 21, 2026Updated 4 months ago
- [ICML 2025] Retraining-Free Merging of Sparse MoE via Hierarchical Clustering☆25Oct 26, 2025Updated 9 months ago
- QuTLASS: CUTLASS-Powered Quantized BLAS for Deep Learning☆199Jul 20, 2026Updated 3 weeks ago
- Official Implementation (Pytorch) of the "Representation Shift: Unifying Token Compression with FlashAttention", ICCV 2025☆36Feb 22, 2026Updated 5 months ago
- PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation (ECCV 2026)☆69Jun 20, 2026Updated last month
- [NeurIPS 2024 Oral🔥] DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs.☆187Apr 24, 2026Updated 3 months ago
- An Efficient and Versatile Inference Engine for Distributed LLM Serving☆66Updated this week
- [ICLR'25] ARB-LLM: Alternating Refined Binarizations for Large Language Models☆31Aug 5, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICML 2026] Elastic Diffusion Transformer: Accelerating SOTA generation models (e.g., Qwen-Image, Hunyuan3d ) through adaptive computatio…☆49May 1, 2026Updated 3 months ago
- [NeurIPS2024] Tune your restoration model with one 3090 GPU!☆90Jan 13, 2025Updated last year
- ☆35Mar 28, 2025Updated last year
- 🧂 [ECCV 2026] Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation☆17Apr 6, 2026Updated 4 months ago
- A Novel Linear Array Pushbroom (LAP) Image Restoration Method. (Accepted by AAAI 2024)☆12Jan 17, 2024Updated 2 years ago
- This repo contains the code for studying the interplay between quantization and sparsity methods☆26Feb 26, 2025Updated last year
- Official repository for paper: O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning☆100Feb 21, 2025Updated last year
- Official PyTorch implementation of CD-MOE☆12Mar 18, 2026Updated 5 months ago
- [ICML 2025] This is the official PyTorch implementation of "ZipAR: Accelerating Auto-regressive Image Generation through Spatial Locality…☆52Mar 25, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Pytorch、Numpy实现NMS、Soft-NMS代码☆12Mar 22, 2021Updated 5 years ago
- [ICLR 2026] Learning to Parallel: Accelerating Diffusion Large Language Models via Learnable Parallel Decoding☆34Jan 27, 2026Updated 6 months ago
- super-resolution; post-training quantization; model compression☆14Nov 10, 2023Updated 2 years ago
- [PR 2024] HTQ: Exploring the High-Dimensional Trade-Off of Mixed-Precision Quantization☆12Jul 16, 2024Updated 2 years ago
- a simple API to use CUPTI☆10Aug 19, 2025Updated 11 months ago
- Code to reproduce the experiments of the ICLR24-paper: "Sparse Model Soups: A Recipe for Improved Pruning via Model Averaging"☆12Oct 14, 2025Updated 10 months ago
- [CVPR 2026 Highlight] Official implementation of Log-linear Sparse Attention (LLSA).☆93May 1, 2026Updated 3 months ago
- The code of paper "O-Mamba: O-shape State-Space Model for Underwater Image Enhancement"☆14Oct 18, 2024Updated last year
- Code for Paper 'DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards'☆18May 21, 2026Updated 2 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Efficient non-uniform quantization with GPTQ for GGUF☆64Sep 17, 2025Updated 11 months ago
- Official Pytorch Implementation of "Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity"☆82Jul 7, 2025Updated last year
- [ICCV-2023] EMQ: Evolving Training-free Proxies for Automated Mixed Precision Quantization☆29Dec 6, 2023Updated 2 years ago
- [NeurIPS 2024] ENAT: Rethinking Spatial-temporal Interactions in Token-based Image Synthesis☆25Nov 28, 2024Updated last year
- image demoireing, moire synthesis☆17Apr 25, 2024Updated 2 years ago
- [ACM MM2025]: MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization☆44Aug 13, 2025Updated last year
- We present Global Search Optics (GSO) to automatically design compact computational imaging systems.☆15Mar 19, 2025Updated last year