Official implementation of LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents
☆26Feb 1, 2026Updated 6 months ago
Alternatives and similar repositories for LRAgent
Users that are interested in LRAgent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2025] Official implementation of "Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning"☆32Oct 20, 2025Updated 9 months ago
- Accurate and fast KV cache compression with a gating mechanism☆28Jul 27, 2026Updated last week
- [ICLR2025] Code and data for paper: Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasonin…☆45Mar 10, 2025Updated last year
- [ACL Findings 2026] Official Implementation of "FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acc…☆32Apr 14, 2026Updated 3 months ago
- [ICML 2024] Official Implementation of SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks☆42Feb 4, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Website for TREC RAG☆14Updated this week
- [NeurIPS 2025] Multipole Attention for Efficient Long Context Reasoning☆24Dec 5, 2025Updated 7 months ago
- Official implementation of LittleBit (NeurIPS 2025) and its follow-up LittleBit-2 (ICML 2026)☆29May 6, 2026Updated 2 months ago
- 😎 Awesome papers on token redundancy reduction☆14Mar 12, 2025Updated last year
- Official codebase for AdaRank: Adaptive Rank Pruning for Enhanced Model Merging (ICLR 2026)☆20Jan 26, 2026Updated 6 months ago
- [NeurIPS 2023] Token-Scaled Logit Distillation for Ternary Weight Generative Language Models☆18Dec 6, 2023Updated 2 years ago
- [ICML 2025] Official PyTorch implementation of "NegMerge: Sign-Consensual Weight Merging for Machine Unlearning"☆16Nov 25, 2025Updated 8 months ago
- Official implementation of "Modeling Multi-Task Model Merging as Adaptive Projective Gradient Descent".☆23May 23, 2025Updated last year
- Triton kernels for Flux☆23Jul 7, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆11May 24, 2024Updated 2 years ago
- PyTorch implementation for "Generative Modeling on Manifolds Through Mixture of Riemannian Diffusion Processes" (ICML 2024).☆13Jul 21, 2024Updated 2 years ago
- [ICLR 2026] AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models☆28Mar 3, 2026Updated 5 months ago
- LLM Inference with Microscaling Format☆35Nov 12, 2024Updated last year
- ☆35Nov 18, 2025Updated 8 months ago
- [PACT'24] GraNNDis. A fast and unified distributed graph neural network (GNN) training framework for both full-batch (full-graph) and min…☆10Aug 13, 2024Updated last year
- To mitigate position bias in LLMs, especially in long-context scenarios, we scale only one dimension of LLMs, reducing position bias and …☆12Jun 18, 2024Updated 2 years ago
- ☆30Oct 2, 2025Updated 10 months ago
- [NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3–4× reduction in memory and 2× decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)☆225Feb 11, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆34May 30, 2025Updated last year
- Pytorch implementation of "Oscillation-Reduced MXFP4 Training for Vision Transformers" on DeiT Model Pre-training☆40May 4, 2026Updated 2 months ago
- PiKV: KV Cache Management System for Mixture of Experts [Efficient ML System]☆62Jul 24, 2026Updated last week
- QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference☆123Mar 6, 2024Updated 2 years ago
- ☆12Oct 9, 2023Updated 2 years ago
- ☆14Oct 3, 2024Updated last year
- [CVPR 2025] Efficient Personalization of Quantized Diffusion Model without Backpropagation☆17Mar 31, 2025Updated last year
- The set of AI agent model implementations, benchmarks, and others used in our paper "The Cost of Dynamic Reasoning: Demystifying AI Agent…☆43Mar 26, 2026Updated 4 months ago
- [NeurIPS 2025] Scaling Speculative Decoding with Lookahead Reasoning☆69Oct 31, 2025Updated 9 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [TrimKV] Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs - [DBTrimKV] Make Each Token Count: Towards Improving Lo…☆17Jul 26, 2026Updated last week
- Code for "Thinking Forward: Memory-Efficient Federated Finetuning of Language Models" (NeurIPS 2024). Spry is a federated learning al…☆13Oct 8, 2024Updated last year
- Contrastive Continual Learning with Importance Sampling and Prototype-Instance Relation Distillation☆12Jul 22, 2024Updated 2 years ago
- [ICML 2024] Sparse Model Inversion: Efficient Inversion of Vision Transformers with Less Hallucination☆14Apr 29, 2025Updated last year
- [CVPR 2022] DiSparse: Disentangled Sparsification for Multitask Model Compression☆13Sep 6, 2022Updated 3 years ago
- Code for paper: Unraveling the Shift of Visual Information Flow in MLLMs: From Phased Interaction to Efficient Inference☆14Jun 7, 2025Updated last year
- An algorithm for weight-activation quantization (W4A4, W4A8) of LLMs, supporting both static and dynamic quantization☆176Nov 26, 2025Updated 8 months ago