DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
☆25May 21, 2026Updated 2 months ago
Alternatives and similar repositories for DUET-VLM
Users that are interested in DUET-VLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding☆29Jun 10, 2026Updated 2 months ago
- [NeurIPS 2025] AutoPrune, a general pruning method for LLM/VLM/VLA☆20Oct 7, 2025Updated 10 months ago
- ☆15Apr 15, 2026Updated 3 months ago
- [NeurIPS 2025] Official repository for “FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language Models”☆33Dec 9, 2025Updated 8 months ago
- [CVPR 2025] DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models☆86Apr 16, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Code for "Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models"☆17Feb 16, 2026Updated 5 months ago
- Official implementation for TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos☆17Jun 9, 2026Updated 2 months ago
- Dataset Distillation via Vision-Language Category Prototype (ICCV 2025)☆17Mar 20, 2026Updated 4 months ago
- "DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation Trainer" [NeurIPS 2025 Accepted]☆19May 22, 2025Updated last year
- WorldCache: Content-Aware Caching for Accelerated Video World Models☆23Jun 28, 2026Updated last month
- ☆16Sep 29, 2024Updated last year
- A collection of VLMs papers, blogs, and projects, with a focus on VLMs in Autonomous Driving and related reasoning techniques.☆11Nov 16, 2024Updated last year
- Code for "Are “Hierarchical” Visual Representations Hierarchical?" in NeurIPS Workshop for Symmetry and Geometry in Neural Representation…☆23Nov 8, 2023Updated 2 years ago
- Official implementation of "ApET: Approximation-Error Guided Token Compression for Efficient VLMs" (CVPR 2026)☆29Jun 29, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official code for **Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity** (PruneSI…☆13Mar 25, 2026Updated 4 months ago
- ☆14May 9, 2023Updated 3 years ago
- ☆35Apr 16, 2026Updated 3 months ago
- Official code repo of PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs☆26Jan 14, 2025Updated last year
- Towards Efficient Multimodal Large Language Models: A Survey on Token Compression☆213Updated this week
- [ACM MM'23] Official implementation of paper "Avatar Knowledge Distillation: Self-ensemble Teacher Paradigm with Uncertainty".☆14Nov 22, 2023Updated 2 years ago
- [ECCV 2026] StAR: Segment Anything Reasoner☆25Apr 2, 2026Updated 4 months ago
- [ICLR 2026] Multi-Head Low-Rank Attention☆34Apr 16, 2026Updated 3 months ago
- Official Implementation of the ECCV 2022 Paper "Class-Incremental Learning with Cross-Space Clustering and Controlled Transfer"☆18Aug 14, 2022Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official Repository for the ICML 2023 paper "BiRT: Bio-inspired Replay in Vision Transformers for Continual Learning"☆16Oct 11, 2023Updated 2 years ago
- [EMNLP 2025 main 🔥] Code for "Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More"☆121Oct 12, 2025Updated 9 months ago
- [CVPR'26 Findings] Source code for "RADSeg Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglom…☆60May 31, 2026Updated 2 months ago
- [WACV 2026] ZonUI-3B — A lightweight, resolution-aware GUI grounding model trained with only 24K samples on a single RTX 4090.☆26Jan 2, 2026Updated 7 months ago
- ☆23Mar 31, 2026Updated 4 months ago
- ☆16Nov 25, 2021Updated 4 years ago
- [ICML2022] "Identity-Disentangled Adversarial Augmentation for Self-Supervised Learning"☆10Jul 24, 2022Updated 4 years ago
- ☆15Oct 21, 2023Updated 2 years ago
- Simple Contourlet Transform in Python☆17Jul 14, 2021Updated 5 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- HWFI: Hybrid Warping Fusion for Video Frame Interpolation. IJCV 2022☆11Sep 7, 2022Updated 3 years ago
- Official implementation of the paper "Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding"☆18Sep 30, 2025Updated 10 months ago
- ☆21Sep 17, 2022Updated 3 years ago
- A lightweight Text-to-Image Retrieval model [Web App]☆29Dec 6, 2024Updated last year
- A Practical Zoom-in GUI Grounding and Behavior-Based Evaluation method.☆26Dec 8, 2025Updated 8 months ago
- Implementation of Latent Replay, a Continual Learning strategy for Real-Time / On The Edge applications☆14May 7, 2020Updated 6 years ago
- [ICCV2025] Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain Reasoning☆25Nov 13, 2025Updated 8 months ago