Official implementation of LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents
☆32Feb 1, 2026Updated 7 months ago
Alternatives and similar repositories for LRAgent
Users that are interested in LRAgent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2025] Official implementation of "Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning"☆32Oct 20, 2025Updated 10 months ago
- Accurate and fast KV cache compression with a gating mechanism☆29Jul 27, 2026Updated last month
- [ICLR2025] Code and data for paper: Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasonin…☆46Mar 10, 2025Updated last year
- [ACL Findings 2026] Official Implementation of "FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acc…☆33Apr 14, 2026Updated 4 months ago
- [ICML 2024] Official Implementation of SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks☆43Feb 4, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Official implementation of LittleBit (NeurIPS 2025) and its follow-up LittleBit-2 (ICML 2026)☆31May 6, 2026Updated 4 months ago
- 😎 Awesome papers on token redundancy reduction☆14Mar 12, 2025Updated last year
- Official codebase for AdaRank: Adaptive Rank Pruning for Enhanced Model Merging (ICLR 2026)☆20Jan 26, 2026Updated 7 months ago
- [NeurIPS 2023] Token-Scaled Logit Distillation for Ternary Weight Generative Language Models☆18Dec 6, 2023Updated 2 years ago
- [ICML 2025] Official PyTorch implementation of "NegMerge: Sign-Consensual Weight Merging for Machine Unlearning"☆16Nov 25, 2025Updated 9 months ago
- The official project website of "SliderQuant: Accurate Post-Training Quantization for LLMs" (accepted to ICLR 2026).☆26Jun 15, 2026Updated 2 months ago
- Official implementation of "Modeling Multi-Task Model Merging as Adaptive Projective Gradient Descent".☆23May 23, 2025Updated last year
- Triton kernels for Flux☆23Jul 7, 2025Updated last year
- [ACL 2026 Main] Code for the paper "ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs"☆33Jun 1, 2026Updated 3 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆11May 24, 2024Updated 2 years ago
- PyTorch implementation for "Generative Modeling on Manifolds Through Mixture of Riemannian Diffusion Processes" (ICML 2024).☆13Jul 21, 2024Updated 2 years ago
- [ICLR 2026] AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models☆29Aug 11, 2026Updated last month
- LLM Inference with Microscaling Format☆35Nov 12, 2024Updated last year
- ☆36Nov 18, 2025Updated 9 months ago
- [PACT'24] GraNNDis. A fast and unified distributed graph neural network (GNN) training framework for both full-batch (full-graph) and min…☆10Aug 13, 2024Updated 2 years ago
- To mitigate position bias in LLMs, especially in long-context scenarios, we scale only one dimension of LLMs, reducing position bias and …☆12Jun 18, 2024Updated 2 years ago
- ☆30Oct 2, 2025Updated 11 months ago
- [NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3–4× reduction in memory and 2× decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)☆225Feb 11, 2026Updated 7 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Pytorch implementation of "Oscillation-Reduced MXFP4 Training for Vision Transformers" on DeiT Model Pre-training☆41May 4, 2026Updated 4 months ago
- PiKV: KV Cache Management System for Mixture of Experts [Efficient ML System]☆63Aug 17, 2026Updated 3 weeks ago
- QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference☆123Mar 6, 2024Updated 2 years ago
- ☆12Oct 9, 2023Updated 2 years ago
- [CVPR 2025] Efficient Personalization of Quantized Diffusion Model without Backpropagation☆17Mar 31, 2025Updated last year
- The set of AI agent model implementations, benchmarks, and others used in our paper "The Cost of Dynamic Reasoning: Demystifying AI Agent…☆46Mar 26, 2026Updated 5 months ago
- [ICML 2024] Sparse Model Inversion: Efficient Inversion of Vision Transformers with Less Hallucination☆14Apr 29, 2025Updated last year
- [CVPR 2022] DiSparse: Disentangled Sparsification for Multitask Model Compression☆13Sep 6, 2022Updated 4 years ago
- Code for paper: Unraveling the Shift of Visual Information Flow in MLLMs: From Phased Interaction to Efficient Inference☆14Jun 7, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Federated Conformal Prediction with Quantile-of-Quantiles (FedCP-QQ)☆11May 6, 2026Updated 4 months ago
- An algorithm for weight-activation quantization (W4A4, W4A8) of LLMs, supporting both static and dynamic quantization☆177Nov 26, 2025Updated 9 months ago
- [TrimKV] Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs - [DBTrimKV] Make Each Token Count: Towards Improving Lo…☆21Jul 26, 2026Updated last month
- [NeurIPS'25] KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems☆18Nov 1, 2025Updated 10 months ago
- PFLoRA-lib: Personalized Federated Learning with LoRA Algorithm Library focusing on privacy-protection, federated-learning, Citation, Ext…☆14Sep 19, 2024Updated last year
- Structured Pruning Adapters in PyTorch☆19Aug 30, 2023Updated 3 years ago
- ☆235Jun 11, 2024Updated 2 years ago