DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
☆25May 21, 2026Updated 3 months ago
Alternatives and similar repositories for DUET-VLM
Users that are interested in DUET-VLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding☆30Jun 10, 2026Updated 2 months ago
- [NeurIPS 2025] AutoPrune, a general pruning method for LLM/VLM/VLA☆20Oct 7, 2025Updated 10 months ago
- ☆16Apr 15, 2026Updated 4 months ago
- [NeurIPS 2025] Official repository for “FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language Models”☆33Dec 9, 2025Updated 8 months ago
- [CVPR 2025] DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models☆86Apr 16, 2026Updated 4 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Code for "Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models"☆18Feb 16, 2026Updated 6 months ago
- Official implementation for TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos☆17Jun 9, 2026Updated 2 months ago
- "DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation Trainer" [NeurIPS 2025 Accepted]☆19May 22, 2025Updated last year
- WorldCache: Content-Aware Caching for Accelerated Video World Models☆23Jun 28, 2026Updated 2 months ago
- (ICLR 2025 Spotlight) Official code repository for Interleaved Scene Graph.☆31Aug 7, 2025Updated last year
- ☆16Sep 29, 2024Updated last year
- A collection of VLMs papers, blogs, and projects, with a focus on VLMs in Autonomous Driving and related reasoning techniques.☆11Nov 16, 2024Updated last year
- Code for "Are “Hierarchical” Visual Representations Hierarchical?" in NeurIPS Workshop for Symmetry and Geometry in Neural Representation…☆23Nov 8, 2023Updated 2 years ago
- 一个基于Langchain和Flask构建的AI大模型自动批改英语作文系统☆23Apr 18, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official implementation of "ApET: Approximation-Error Guided Token Compression for Efficient VLMs" (CVPR 2026)☆30Jun 29, 2026Updated 2 months ago
- ☆36Apr 16, 2026Updated 4 months ago
- [ICML25] Agentic Compression Benchmark (ACBench)☆19Jul 2, 2025Updated last year
- Official code repo of PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs☆26Jan 14, 2025Updated last year
- Towards Efficient Multimodal Large Language Models: A Survey on Token Compression☆221Aug 10, 2026Updated 2 weeks ago
- 对youtu的training_free_grpo的测试以及修改☆24Nov 6, 2025Updated 9 months ago
- [ECCV 2026] StAR: Segment Anything Reasoner☆25Apr 2, 2026Updated 4 months ago
- ☆13Jul 1, 2024Updated 2 years ago
- [ICLR 2026] Multi-Head Low-Rank Attention☆33Apr 16, 2026Updated 4 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Official Repository for the ICML 2023 paper "BiRT: Bio-inspired Replay in Vision Transformers for Continual Learning"☆16Oct 11, 2023Updated 2 years ago
- [EMNLP 2025 main 🔥] Code for "Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More"☆121Oct 12, 2025Updated 10 months ago
- [CVPR'26 Findings] Source code for "RADSeg Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglom…☆61May 31, 2026Updated 2 months ago
- [WACV 2026] ZonUI-3B — A lightweight, resolution-aware GUI grounding model trained with only 24K samples on a single RTX 4090.☆26Jan 2, 2026Updated 7 months ago
- 3D Sign Language Production for ASL☆17Dec 4, 2020Updated 5 years ago
- ☆16Nov 25, 2021Updated 4 years ago
- ☆15Oct 21, 2023Updated 2 years ago
- CVPR25☆28Jul 2, 2025Updated last year
- HWFI: Hybrid Warping Fusion for Video Frame Interpolation. IJCV 2022☆11Sep 7, 2022Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Official implementation of the paper "Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding"☆18Sep 30, 2025Updated 11 months ago
- ☆21Sep 17, 2022Updated 3 years ago
- A simple visual test-time scaling method for GUI agent grounding☆26Dec 7, 2025Updated 8 months ago
- Code for paper: "What’s in the Image? A Deep-Dive into the Vision of Vision Language Models" (CVPR 2025)☆18May 1, 2025Updated last year
- A Practical Zoom-in GUI Grounding and Behavior-Based Evaluation method.☆28Updated this week
- Contrastive self-supervised learning using Rényi divergence☆14Oct 21, 2022Updated 3 years ago
- An arbitrage bot is a smart contract connected to an external automation script that controls its operation.☆2,690Updated this week