☆73Jul 17, 2024Updated 2 years ago
Alternatives and similar repositories for QA-ViT
Users that are interested in QA-ViT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆24Apr 29, 2025Updated last year
- ☆13May 21, 2024Updated 2 years ago
- FuseCap: Leveraging Large Language Models for Enriched Fused Image Captions☆55Apr 17, 2024Updated 2 years ago
- Multimodal Semi-Supervised Learning for Text Recognition (SemiMTR)☆83Sep 12, 2023Updated 3 years ago
- Ada-LISTA: Learned Solvers Adaptive to Varying Models☆11Feb 18, 2020Updated 6 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- EMMA [TMLR 2025]☆14Sep 25, 2025Updated last year
- ☆13Mar 8, 2025Updated last year
- Code for paper DMCVR: Morphology-Guided Diffusion Model for 3D Cardiac Volume Reconstruction☆14Jan 12, 2024Updated 2 years ago
- official repo for paper "[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs"☆22Apr 23, 2025Updated last year
- STVQA and TextVQA OCR results from Amazon Text in Image pipeline☆12Jul 18, 2022Updated 4 years ago
- ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs☆30Aug 15, 2025Updated last year
- ☆16Dec 25, 2021Updated 4 years ago
- ☆14May 3, 2022Updated 4 years ago
- ☆32Jul 29, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- This repository contains code for AAAI2025 paper "Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal …☆26Aug 18, 2025Updated last year
- Code for paper: "Privately generating tabular data using language models".☆16Jun 13, 2023Updated 3 years ago
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities (ICML 2024)☆334Jan 20, 2025Updated last year
- [CVPR'24 Highlight] Implementation of "Causal-CoG: A Causal-Effect Look at Context Generation for Boosting Multi-modal Language Models"☆17Sep 12, 2024Updated 2 years ago
- Some papers about *diverse* image (a few videos) captioning☆25Apr 4, 2023Updated 3 years ago
- [CVPR 2024] Do you remember? Dense Video Captioning with Cross-Modal Memory Retrieval☆67Jun 19, 2024Updated 2 years ago
- "Towards Improving Document Understanding: An Exploration on Text-Grounding via MLLMs" 2023☆16Nov 28, 2024Updated last year
- ☆22Jun 5, 2025Updated last year
- ☆17Aug 11, 2023Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering [ACM MM'24]☆10Jul 22, 2024Updated 2 years ago
- Code for Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models☆91Jun 28, 2024Updated 2 years ago
- Code for "AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs"☆26Oct 9, 2025Updated 11 months ago
- ☆22Apr 27, 2024Updated 2 years ago
- ☆28Jul 18, 2025Updated last year
- Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types☆32Jul 16, 2025Updated last year
- ☆23Sep 28, 2023Updated 2 years ago
- The repo for "On-the-fly Modulation for Balanced Multimodal Learning", T-PAMI 2024☆21Sep 29, 2024Updated last year
- Official Repository for CVPR 2022 paper "REX: Reasoning-aware and Grounded Explanation"☆22Nov 21, 2023Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Code for "Don't trust your eyes: on the (un)reliability of feature visualizations" (ICML 2024)☆33Nov 15, 2023Updated 2 years ago
- MUSIC-AVQA, CVPR2022 (ORAL)☆100Dec 30, 2022Updated 3 years ago
- Official Repository for "Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality" (ECCV 2024)☆16Oct 29, 2024Updated last year
- Tempo: Small Vision-Language Models are Smart Compressors for Long Video Understanding, ECCV 2026☆85Jul 26, 2026Updated last month
- Code for Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense? [COLM 2024]☆24Aug 13, 2024Updated 2 years ago
- ☆11May 24, 2024Updated 2 years ago
- ☆12Dec 15, 2023Updated 2 years ago