Cross-Self KV Cache Pruning for Efficient Vision-Language Inference
β10Dec 15, 2024Updated last year
Alternatives and similar repositories for CSP
Users that are interested in CSP are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- (ICME24) This is the offical repository of iDAT: inverse Distillation Adapter-Tuning.β13Apr 3, 2024Updated 2 years ago
- [EMNLP 2024 Findings π₯] Official implementation of ": LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inβ¦β103Nov 9, 2024Updated last year
- ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Modelsβ16Sep 27, 2024Updated last year
- [NAACL 2025π₯] MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inferenceβ22Jun 19, 2025Updated last year
- Official Implementation of CODEβ17Sep 26, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ECCV 2024] Efficient Inference of Vision Instruction-Following Models with Elastic Cacheβ43Jul 26, 2024Updated last year
- This is the open-source code for TokenCarve.β25Jan 23, 2026Updated 5 months ago
- [ICLR2026] Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perceptionβ17Jan 26, 2026Updated 5 months ago
- ICLR2024: Neural Architecture Retrievalβ16Mar 13, 2024Updated 2 years ago
- Official resource for paper Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models (ACL 20β¦β18Aug 12, 2024Updated last year
- [NeurIPS 2024] Official Repository of Multi-Object Hallucination in Vision-Language Modelsβ37Nov 13, 2024Updated last year
- (CVPR 2025) PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reductionβ151Mar 6, 2025Updated last year
- Official PyTorch Implementation for the "What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modβ¦β20Sep 26, 2024Updated last year
- CVPR2024 highlight.β13Oct 10, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [NeurIPS 2024] TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent Collaborationβ25Oct 17, 2024Updated last year
- (ACM MM24) This is the offical repository of GIST: Improving Parameter Efficient Fine Tuning via Knowledge Interaction.β11Jan 28, 2024Updated 2 years ago
- β48May 9, 2026Updated 2 months ago
- Official implementation of paper "Masked Distillation with Receptive Tokens", ICLR 2023.β10Mar 13, 2023Updated 3 years ago
- The official implementation of the paper **LVChat: Facilitating Long Video Comprehension**β14Apr 15, 2024Updated 2 years ago
- [CVPR 2025] Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attentionβ68Jul 16, 2024Updated 2 years ago
- Official PyTorch implementation of Agglomerative Token Clustering presented at ECCV 2024β20Sep 19, 2024Updated last year
- β14Apr 11, 2022Updated 4 years ago
- Benchmarks for Macro Neural Architecture Search; used and described in the paper "Local Search is a Remarkably Strong Baseline for Neuralβ¦β13Jul 25, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- CAD - Memory Efficient Convolutional Adapter for Segment Anythingβ12Oct 4, 2024Updated last year
- DIRT:Deep Learning Enhanced Item Response Theory for Cognitive Diagnosisβ25Dec 4, 2019Updated 6 years ago
- HallE-Control: Controlling Object Hallucination in LMMsβ32Apr 10, 2024Updated 2 years ago
- Pytorch implementation of our paper accepted by ICML 2023 -- "Bi-directional Masks for Efficient N:M Sparse Training"β13Jun 7, 2023Updated 3 years ago
- [ACM MM'23] Official implementation of paper "Avatar Knowledge Distillation: Self-ensemble Teacher Paradigm with Uncertainty".β14Nov 22, 2023Updated 2 years ago
- Latex Template for Northwestern Polytechnical University(NWPU) Reportβ17Oct 16, 2023Updated 2 years ago
- The code for "AttentionPredictor: Temporal Pattern Matters for Efficient LLM Inference", Qingyue Yang, Jie Wang, Xing Li, Zhihai Wang, Chβ¦β29Jul 15, 2025Updated last year
- Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Modelβ36Jan 8, 2025Updated last year
- Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Trainingβ16Jul 1, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- U08M11002 δΏ‘ε·δΈη³»η»οΌθ₯Ώεε·₯δΈε€§ε¦β15Jan 19, 2024Updated 2 years ago
- EUV Layer Hotspot Detection Benchmark Suitβ20Mar 8, 2021Updated 5 years ago
- β17Aug 8, 2024Updated last year
- Official InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Showsβ20Nov 4, 2025Updated 8 months ago
- Code for "Dual Focal Loss for Calibration" (ICML 2023)β31Apr 21, 2025Updated last year
- Pytorch implementation of TSE attentionβ16Jul 9, 2021Updated 5 years ago
- Source code of paper ''KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing''β31Oct 24, 2024Updated last year