[NeurIPS 2024] Official Code for the Paper "Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning"
☆27Apr 8, 2025Updated last year
Alternatives and similar repositories for MultiModal-Task-Vector
Users that are interested in MultiModal-Task-Vector are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official repository for Efficient Multi-modal Long Context Learning for Training-free Adaptation (ICML 2025)☆20Dec 27, 2025Updated 8 months ago
- AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence☆11Mar 2, 2025Updated last year
- 【NeurIPS 2024】The implementation of LIVE: Learnable In-Context Vector for Visual Question Answering https://arxiv.org/abs/2406.13185☆23May 31, 2025Updated last year
- Official Codebase for "Generative Multimodal Model Features Are Discriminative Vision-Language Classifiers"☆29Jun 7, 2025Updated last year
- Function Vectors in Large Language Models (ICLR 2024)☆200Apr 30, 2026Updated 4 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆31Feb 10, 2025Updated last year
- [ICCV 2025] Official code for "AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning"☆65Oct 9, 2025Updated 11 months ago
- Code for ICLR 2025 Paper: Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs☆25May 7, 2025Updated last year
- PyTorch code for "Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training"☆39Mar 4, 2024Updated 2 years ago
- Official PyTorch Implementation for Vision-Language Models Create Cross-Modal Task Representations, ICML 2025☆35May 1, 2025Updated last year
- Frequently updated list of dLLM (Diffusion Large Language Models) papers, models, and other resources☆55Sep 1, 2026Updated 2 weeks ago
- [CVPR 2024] Official Code for the Paper "Compositional Chain-of-Thought Prompting for Large Multimodal Models"☆144Jun 20, 2024Updated 2 years ago
- Implementation and dataset for paper "Can MLLMs Perform Text-to-Image In-Context Learning?"☆48Jun 2, 2025Updated last year
- ☆33May 9, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Official implementation of the CVPR '25 highlight paper "Compositional Caching for Training-free Open-vocabulary Attribute Detection"☆23Dec 23, 2024Updated last year
- ☆12Jan 10, 2025Updated last year
- [WACV 26] Official code for the paper Safe Vision-Language Models via Unsafe Weights Manipulation☆16Mar 3, 2026Updated 6 months ago
- ☆28Jul 18, 2025Updated last year
- A library of visualization tools for the interpretability and hallucination analysis of large vision-language models (LVLMs).☆43May 22, 2025Updated last year
- ☆22Jan 7, 2025Updated last year
- A curated lists of self-taught materials including research blogs☆16Dec 12, 2016Updated 9 years ago
- [CVPR 2025] Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Att…☆84Oct 9, 2025Updated 11 months ago
- ☆14Oct 25, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- MinPrompt: Graph-based Minimal Prompt Data Augmentation for Few-shot Question Answering☆14May 3, 2024Updated 2 years ago
- [ICML 2025] Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMM☆20May 22, 2025Updated last year
- ☆12Jun 29, 2024Updated 2 years ago
- Code for paper: Unraveling the Shift of Visual Information Flow in MLLMs: From Phased Interaction to Efficient Inference☆14Jun 7, 2025Updated last year
- ☆15Dec 11, 2024Updated last year
- ☆11Jan 14, 2021Updated 5 years ago
- Building Llama 3 from scratch using PyTorch☆13Sep 1, 2024Updated 2 years ago
- AdaICL: Which Examples to Annotate of In-Context Learning? Towards Effective and Efficient Selection☆19Oct 30, 2023Updated 2 years ago
- [CVPR Findings 2026] Large Multimodal Models as General In-Context Classifiers☆26Mar 1, 2026Updated 6 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Code for In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering☆202Feb 13, 2025Updated last year
- CCF 2021 BDCI 千言-问题匹配鲁棒性评测 A榜 rank 29th, B榜 rank 15th☆14Jan 5, 2022Updated 4 years ago
- Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering [ACM MM'24]☆10Jul 22, 2024Updated 2 years ago
- ☆11Apr 29, 2020Updated 6 years ago
- The Code for Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models☆19Oct 4, 2024Updated last year
- From-Classification-to-Clinical☆13Apr 26, 2024Updated 2 years ago
- [Paper] Code for the EMNLP2023 (Findings) paper "Global Structure Knowledge-Guided Relation Extraction Method for Visually-Rich Document"☆17Dec 1, 2023Updated 2 years ago