[IJCAI 2025] Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
☆36Nov 25, 2025Updated 7 months ago
Alternatives and similar repositories for awesome-captioning-evaluation
Users that are interested in awesome-captioning-evaluation are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17Feb 20, 2025Updated last year
- [CVPR 2023 & IJCV 2025] Positive-Augmented Contrastive Learning for Image and Video Captioning Evaluation☆66Jul 29, 2025Updated 11 months ago
- This is the official repository for the paper "Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction". ICCV …☆27May 13, 2026Updated 2 months ago
- Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models. ECCV 2024☆68Aug 10, 2024Updated last year
- [ECCV'24] Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities☆52Jul 2, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICLR 2026] "Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals"☆49Mar 6, 2026Updated 4 months ago
- This repository contains a curated list of research papers and resources focusing on saliency and scanpath prediction, human attention, h…☆66May 9, 2025Updated last year
- [ACL 2024] FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model☆17Apr 28, 2025Updated last year
- [ICCV 2025] MissRAG: Addressing the Missing Modality Challenge in Multimodal Large Language Models☆26May 12, 2026Updated 2 months ago
- ☆10Sep 2, 2021Updated 4 years ago
- Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training☆16Jul 1, 2025Updated last year
- 我在校园的各项API,自动运行脚本,支持多人☆12Jun 28, 2022Updated 4 years ago
- HIPPO 🦛 is an explainable AI method and toolkit for weakly-supervised models in computational pathology. It enables hypothesis testing o…☆19Jan 15, 2025Updated last year
- ☆13Dec 12, 2022Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- CLIPScore EMNLP code☆251Dec 16, 2022Updated 3 years ago
- [NeurIPS 2023] A faithful benchmark for vision-language compositionality☆93Feb 13, 2024Updated 2 years ago
- [CVPR 2024 Highlight] Official repository of the paper "The devil is in the fine-grained details: Evaluating open-vocabulary object detec…☆68Apr 4, 2025Updated last year
- Microsoft question-answering dataset☆10Jun 16, 2023Updated 3 years ago
- [ICLR 2025] - Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion☆70Nov 30, 2025Updated 7 months ago
- The code for the Ensemble everything everywhere: Multi-scale aggregation for adversarial robustness paper☆23Nov 8, 2024Updated last year
- Web Service for Human Pose Annotation☆10Oct 21, 2018Updated 7 years ago
- [CVPR24 Highlights] Polos: Multimodal Metric Learning from Human Feedback for Image Captioning☆33Jun 12, 2026Updated last month
- PyTorch code for BMVC 2019 paper: Embodied Vision-and-Language Navigation with Dynamic Convolutional Filters☆20Jan 4, 2023Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [CVPR '26] CaptionQA: Is Your Caption as Useful as the Image Itself?☆38Mar 3, 2026Updated 4 months ago
- PyTorch code for the paper: "Perceive, Transform, and Act: Multi-Modal Attention Networks for Vision-and-Language Navigation"☆19Aug 5, 2021Updated 4 years ago
- [IEEE TMM 2023] This is the official repo of the paper "Perceptual Quality Improvement in Videoconferencing using Keyframes-based GAN".☆17Dec 10, 2024Updated last year
- FeelingBlue: A Corpus for Understanding the Emotional Connotation of Color in Context, accepted at TACL 2022, presented at ACL 2023☆13Dec 28, 2023Updated 2 years ago
- The implementation for "DEER: Descriptive Knowledge Graph for Explaining Entity Relationships" (EMNLP '22)☆11Oct 31, 2022Updated 3 years ago
- Using CNN for classifying 101 different food categories - using VGG16, Alex Net and SVM☆10Jan 6, 2020Updated 6 years ago
- ☆12Jan 16, 2024Updated 2 years ago
- SOMOSPIE (Soil Moisture Spatial Inference Engine) consists of a Jupyter Notebook and a suite of machine learning methods to process input…☆16Sep 16, 2025Updated 10 months ago
- [ICCVW 25] LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning☆160Aug 8, 2025Updated 11 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Umap is a Python library that transforms OpenStreetMap data into customized maps with minimal code. Create minimalist or multi-layered ma…☆15Jul 7, 2026Updated 2 weeks ago
- DJI Phantom 4 Multispectral raw image radiometric calibration☆13Jun 4, 2021Updated 5 years ago
- A currency rate converter App.☆15Sep 5, 2019Updated 6 years ago
- C++, OpenMP and CUDA implementation of Mean Shift clustering algorithm☆14Apr 27, 2020Updated 6 years ago
- ☆49Jun 26, 2026Updated 3 weeks ago
- Presents an optimised Sentinel-1-based Soil Moisture estimation workflow on tilted topography with permanent vegetation cover☆16Aug 18, 2022Updated 3 years ago
- Code for the ACL 2019 paper "Observing Dialogue in Therapy: Categorizing and Forecasting Behavioral Codes"☆14Jun 11, 2022Updated 4 years ago