[NeurIPS 2024] Official Code for the Paper "Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning"
☆27Apr 8, 2025Updated last year
Alternatives and similar repositories for MultiModal-Task-Vector
Users that are interested in MultiModal-Task-Vector are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official repository for Efficient Multi-modal Long Context Learning for Training-free Adaptation (ICML 2025)☆20Dec 27, 2025Updated 9 months ago
- AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence☆11Mar 2, 2025Updated last year
- 【NeurIPS 2024】The implementation of LIVE: Learnable In-Context Vector for Visual Question Answering https://arxiv.org/abs/2406.13185☆23May 31, 2025Updated last year
- Official Codebase for "Generative Multimodal Model Features Are Discriminative Vision-Language Classifiers"☆29Jun 7, 2025Updated last year
- Official Pytorch Implementation of "Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generati…☆12Aug 26, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆11Nov 5, 2024Updated last year
- [ICCV 2025] Official code for "AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning"☆65Oct 9, 2025Updated last year
- Code for ICLR 2025 Paper: Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs☆25May 7, 2025Updated last year
- PyTorch code for "Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training"☆39Mar 4, 2024Updated 2 years ago
- Official PyTorch Implementation for Vision-Language Models Create Cross-Modal Task Representations, ICML 2025☆35May 1, 2025Updated last year
- Frequently updated list of dLLM (Diffusion Large Language Models) papers, models, and other resources☆57Sep 1, 2026Updated last month
- [CVPR 2024] Official Code for the Paper "Compositional Chain-of-Thought Prompting for Large Multimodal Models"☆143Jun 20, 2024Updated 2 years ago
- ☆33May 9, 2025Updated last year
- Official implementation of the CVPR '25 highlight paper "Compositional Caching for Training-free Open-vocabulary Attribute Detection"☆23Dec 23, 2024Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [WACV 26] Official code for the paper Safe Vision-Language Models via Unsafe Weights Manipulation☆16Mar 3, 2026Updated 7 months ago
- Visual interpretation of deep learning model in ECG classification: A comprehensive evaluation of feature attribution methods☆14Jun 9, 2025Updated last year
- Official implementation of the WACV 2025 paper "3D Part Segmentation via Geometric Aggregation of 2D Visual Features"☆25Jun 8, 2025Updated last year
- ☆28Jul 18, 2025Updated last year
- ☆22Jan 7, 2025Updated last year
- A curated lists of self-taught materials including research blogs☆16Dec 12, 2016Updated 9 years ago
- [CVPR 2025] Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Att…☆84Oct 9, 2025Updated last year
- ☆14Oct 25, 2024Updated last year
- [ICML 2025] Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMM☆20May 22, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆12Jun 29, 2024Updated 2 years ago
- Code for paper: Unraveling the Shift of Visual Information Flow in MLLMs: From Phased Interaction to Efficient Inference☆14Jun 7, 2025Updated last year
- Experiments with representation engineering☆14Feb 28, 2024Updated 2 years ago
- ☆12Jan 14, 2021Updated 5 years ago
- Building Llama 3 from scratch using PyTorch☆13Sep 1, 2024Updated 2 years ago
- Towards Robust and Relible Multimodal Misinformation Recognition with Incomplete Modality☆15May 11, 2026Updated 4 months ago
- 基于BERT和MRC框架实现的嵌套命名实体识别☆19Mar 13, 2022Updated 4 years ago
- Official Codes for Fine-Grained Visual Prompting, NeurIPS 2023☆57Feb 1, 2024Updated 2 years ago
- Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering [ACM MM'24]☆10Jul 22, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆11Apr 29, 2020Updated 6 years ago
- The Code for Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models☆19Oct 4, 2024Updated 2 years ago
- From-Classification-to-Clinical☆13Apr 26, 2024Updated 2 years ago
- [Paper] Code for the EMNLP2023 (Findings) paper "Global Structure Knowledge-Guided Relation Extraction Method for Visually-Rich Document"☆17Dec 1, 2023Updated 2 years ago
- The code repository for "OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions"☆13Feb 21, 2025Updated last year
- ALPHA: AnomaLous Physiological Health Assessment Using Large Language Models (AI Health Summit 23)☆19Feb 25, 2025Updated last year
- A javascript program to read data from multiple Polar devices (H10 & Verity Sense)☆17Mar 21, 2025Updated last year