Official implementation of LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents
☆32Sep 24, 2026Updated last week
Alternatives and similar repositories for LRAgent
Users that are interested in LRAgent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2025] Official implementation of "Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning"☆32Oct 20, 2025Updated 11 months ago
- Accurate and fast KV cache compression with a gating mechanism☆31Jul 27, 2026Updated 2 months ago
- [ICLR2025] Code and data for paper: Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasonin…☆46Mar 10, 2025Updated last year
- [ACL Findings 2026] Official Implementation of "FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acc…☆33Apr 14, 2026Updated 5 months ago
- Official PyTorch implementation of Safety-Guided Flow (SGF): "Safety-Guided Flow (SGF): A Unified Framework for Negative Guidance in Safe…☆15Feb 25, 2026Updated 7 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [NeurIPS 2025] Multipole Attention for Efficient Long Context Reasoning☆26Dec 5, 2025Updated 9 months ago
- Official implementation of LittleBit (NeurIPS 2025) and its follow-up LittleBit-2 (ICML 2026)☆32May 6, 2026Updated 4 months ago
- Official codebase for AdaRank: Adaptive Rank Pruning for Enhanced Model Merging (ICLR 2026)☆20Jan 26, 2026Updated 8 months ago
- [ICML 2025] Official PyTorch implementation of "NegMerge: Sign-Consensual Weight Merging for Machine Unlearning"☆16Nov 25, 2025Updated 10 months ago
- The official project website of "SliderQuant: Accurate Post-Training Quantization for LLMs" (accepted to ICLR 2026).☆27Jun 15, 2026Updated 3 months ago
- Official implementation of "Modeling Multi-Task Model Merging as Adaptive Projective Gradient Descent".☆23May 23, 2025Updated last year
- [ACL 2026 Main] Code for the paper "ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs"☆32Sep 21, 2026Updated last week
- ☆11May 24, 2024Updated 2 years ago
- PyTorch implementation for "Generative Modeling on Manifolds Through Mixture of Riemannian Diffusion Processes" (ICML 2024).☆13Jul 21, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- 西安电子科技大学学位论文 Typst 模板(硕士学术 / 专业学位、本科毕业设计论文)。☆16Sep 17, 2026Updated 2 weeks ago
- [ICLR 2026] AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models☆31Aug 11, 2026Updated last month
- LLM Inference with Microscaling Format☆35Nov 12, 2024Updated last year
- ☆36Nov 18, 2025Updated 10 months ago
- [PACT'24] GraNNDis. A fast and unified distributed graph neural network (GNN) training framework for both full-batch (full-graph) and min…☆10Aug 13, 2024Updated 2 years ago
- To mitigate position bias in LLMs, especially in long-context scenarios, we scale only one dimension of LLMs, reducing position bias and …☆12Jun 18, 2024Updated 2 years ago
- ☆30Sep 11, 2026Updated 3 weeks ago
- [NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3–4× reduction in memory and 2× decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)☆225Feb 11, 2026Updated 7 months ago
- ☆35May 30, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Pytorch implementation of "Oscillation-Reduced MXFP4 Training for Vision Transformers" on DeiT Model Pre-training☆41May 4, 2026Updated 4 months ago
- ☆26Oct 9, 2025Updated 11 months ago
- PiKV: KV Cache Management System for Mixture of Experts [Efficient ML System]☆63Aug 17, 2026Updated last month
- QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference☆123Mar 6, 2024Updated 2 years ago
- ☆12Oct 9, 2023Updated 2 years ago
- ☆14Oct 3, 2024Updated 2 years ago
- [CVPR 2025] Efficient Personalization of Quantized Diffusion Model without Backpropagation☆17Mar 31, 2025Updated last year
- ☆10Sep 13, 2022Updated 4 years ago
- The set of AI agent model implementations, benchmarks, and others used in our paper "The Cost of Dynamic Reasoning: Demystifying AI Agent…☆47Mar 26, 2026Updated 6 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [NeurIPS 2025] Scaling Speculative Decoding with Lookahead Reasoning☆69Oct 31, 2025Updated 11 months ago
- Code for "Thinking Forward: Memory-Efficient Federated Finetuning of Language Models" (NeurIPS 2024). Spry is a federated learning al…☆13Oct 8, 2024Updated last year
- [ICML 2024] Sparse Model Inversion: Efficient Inversion of Vision Transformers with Less Hallucination☆15Apr 29, 2025Updated last year
- [CVPR 2022] DiSparse: Disentangled Sparsification for Multitask Model Compression☆14Sep 6, 2022Updated 4 years ago
- Code for paper: Unraveling the Shift of Visual Information Flow in MLLMs: From Phased Interaction to Efficient Inference☆14Jun 7, 2025Updated last year
- Federated Conformal Prediction with Quantile-of-Quantiles (FedCP-QQ)☆11May 6, 2026Updated 4 months ago
- An algorithm for weight-activation quantization (W4A4, W4A8) of LLMs, supporting both static and dynamic quantization☆178Nov 26, 2025Updated 10 months ago