iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models (ICLR2026)
☆23Jun 24, 2026Updated last month
Alternatives and similar repositories for iLLaVA
Users that are interested in iLLaVA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆11Dec 20, 2024Updated last year
- [ICML 2025] Code for "R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts"☆19Mar 10, 2025Updated last year
- A星算法及其可视化实现(python)☆13Dec 4, 2019Updated 6 years ago
- ☆29May 13, 2025Updated last year
- ☆13Feb 2, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- (CVPR 2025) PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction☆151Mar 6, 2025Updated last year
- ☆22Mar 3, 2025Updated last year
- Extending context length of visual language models☆12Dec 18, 2024Updated last year
- How to plot for papers, slides, demos, etc.☆10Apr 7, 2022Updated 4 years ago
- Prompt Tuning on Graph-augmented Low-resource Text Classification. In TKDE 2024.☆15Jan 20, 2025Updated last year
- Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model☆13Feb 11, 2025Updated last year
- PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding☆30Jun 10, 2026Updated 2 months ago
- Code for ACL 2024 paper "Soft Self-Consistency Improves Language Model Agents"☆25Sep 11, 2024Updated last year
- 📷 Python package and CLI utility to create photo mosaics - now with GPU support☆18Mar 6, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆47Nov 8, 2024Updated last year
- ☆11Jul 26, 2024Updated 2 years ago
- [NeurIPS 2023]Federated Learning with Bilateral Curation for Partially Class-Disjoint Data☆14Aug 1, 2025Updated last year
- [ICCV 2025] Official code for "AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning"☆65Oct 9, 2025Updated 10 months ago
- ☆11Sep 30, 2024Updated last year
- CVPR25☆28Jul 2, 2025Updated last year
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆23Jan 25, 2026Updated 6 months ago
- Official PyTorch implementation for the ICML 2023 paper "Out-of-Distribution Generalization of Federated Learning via Implicit Invariant …☆14Oct 31, 2023Updated 2 years ago
- MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer☆50Sep 6, 2024Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [ICME 2024 Oral] DARA: Domain- and Relation-aware Adapters Make Parameter-efficient Tuning for Visual Grounding☆22Feb 26, 2025Updated last year
- ☆18Jul 8, 2025Updated last year
- Official Repository for NeurIPS'25 Paper "Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task"☆23May 18, 2026Updated 3 months ago
- KARL: Knowledge-Aware Reasoning and Reinforcement Learning for Knowledge-Intensive Visual Grounding☆70Apr 5, 2026Updated 4 months ago
- Official repository for Efficient Multi-modal Long Context Learning for Training-free Adaptation (ICML 2025)☆19Dec 27, 2025Updated 7 months ago
- [NeurIPS 2025] Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models☆60Sep 29, 2025Updated 10 months ago
- ☆50Oct 28, 2024Updated last year
- Code for EMNLP25 paper "Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning"☆24Feb 18, 2026Updated 6 months ago
- Synthesizing Efficient Data with Diffusion Models for Person Re-Identification Pre-Training☆11Jan 23, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ERGO (Efficient Reasoning & Guided Observation) is a large vision-language model trained with reinforcement learning on efficiency object…☆19Feb 25, 2026Updated 5 months ago
- ☆10Dec 3, 2024Updated last year
- [ICLR 26] The official code repository for the paper "Mirage or Method? How Model–Task Alignment Induces Divergent RL Conclusions".☆19Feb 9, 2026Updated 6 months ago
- [ICLR 2025] Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models☆63Jan 22, 2025Updated last year
- Code for the paper "Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns"☆18Mar 15, 2024Updated 2 years ago
- Official repo for "Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge" ICLR2025☆112Mar 14, 2025Updated last year
- Matryoshka Multimodal Models☆123Jan 22, 2025Updated last year