[NeurIPS 2024] Official Code for the Paper "Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning"
☆27Apr 8, 2025Updated last year
Alternatives and similar repositories for MultiModal-Task-Vector
Users that are interested in MultiModal-Task-Vector are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official repository for Efficient Multi-modal Long Context Learning for Training-free Adaptation (ICML 2025)☆20Dec 27, 2025Updated 8 months ago
- AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence☆10Mar 2, 2025Updated last year
- 【NeurIPS 2024】The implementation of LIVE: Learnable In-Context Vector for Visual Question Answering https://arxiv.org/abs/2406.13185☆23May 31, 2025Updated last year
- Official Codebase for "Generative Multimodal Model Features Are Discriminative Vision-Language Classifiers"☆27Jun 7, 2025Updated last year
- ☆11Nov 5, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆31Feb 10, 2025Updated last year
- Code for ICLR 2025 Paper: Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs☆25May 7, 2025Updated last year
- PyTorch code for "Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training"☆39Mar 4, 2024Updated 2 years ago
- Official PyTorch Implementation for Vision-Language Models Create Cross-Modal Task Representations, ICML 2025☆34May 1, 2025Updated last year
- Frequently updated list of dLLM (Diffusion Large Language Models) papers, models, and other resources☆55Aug 19, 2026Updated last week
- Implementation and dataset for paper "Can MLLMs Perform Text-to-Image In-Context Learning?"☆48Jun 2, 2025Updated last year
- ☆33May 9, 2025Updated last year
- Official implementation of the CVPR '25 highlight paper "Compositional Caching for Training-free Open-vocabulary Attribute Detection"☆23Dec 23, 2024Updated last year
- ☆12Jan 10, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Official implementation of the WACV 2025 paper "3D Part Segmentation via Geometric Aggregation of 2D Visual Features"☆25Jun 8, 2025Updated last year
- SNARE Dataset with MATCH and LaGOR models☆23Mar 27, 2024Updated 2 years ago
- A latent-variable model for learning bilingual word embedding mappings☆19Feb 11, 2019Updated 7 years ago
- ☆28Jul 18, 2025Updated last year
- WWW 2024: New Frontiers of Knowledge Graph Reasoning: Recent Advances and Future Trends☆18Mar 24, 2026Updated 5 months ago
- weixin125个人健康数据管理系统的设计与实现微信小程序+ssm后端毕业源码案例设计☆10Feb 28, 2024Updated 2 years ago
- A library of visualization tools for the interpretability and hallucination analysis of large vision-language models (LVLMs).☆43May 22, 2025Updated last year
- ☆22Jan 7, 2025Updated last year
- A curated lists of self-taught materials including research blogs☆16Dec 12, 2016Updated 9 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆13Oct 25, 2024Updated last year
- MinPrompt: Graph-based Minimal Prompt Data Augmentation for Few-shot Question Answering☆14May 3, 2024Updated 2 years ago
- Code for paper: Unraveling the Shift of Visual Information Flow in MLLMs: From Phased Interaction to Efficient Inference☆14Jun 7, 2025Updated last year
- Experiments with representation engineering☆14Feb 28, 2024Updated 2 years ago
- ☆11Jan 14, 2021Updated 5 years ago
- Building Llama 3 from scratch using PyTorch☆13Sep 1, 2024Updated last year
- [CVPR Findings 2026] Large Multimodal Models as General In-Context Classifiers☆25Mar 1, 2026Updated 5 months ago
- Official Code Repository for paper "HYDRA: Model Factorization Framework for Black-Box LLM Personalization"☆16Oct 7, 2024Updated last year
- Official Codes for Fine-Grained Visual Prompting, NeurIPS 2023☆56Feb 1, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering [ACM MM'24]☆10Jul 22, 2024Updated 2 years ago
- ☆11Apr 29, 2020Updated 6 years ago
- The Code for Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models☆18Oct 4, 2024Updated last year
- We enable LLM with personalization capability☆11Nov 16, 2023Updated 2 years ago
- [Paper] Code for the EMNLP2023 (Findings) paper "Global Structure Knowledge-Guided Relation Extraction Method for Visually-Rich Document"☆17Dec 1, 2023Updated 2 years ago
- The code repository for "OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions"☆13Feb 21, 2025Updated last year
- An API that detect expiration date from the product package's picture based on Deep Learning Algorithms☆11Jun 4, 2022Updated 4 years ago