[IJCAI 2025] Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
☆37Nov 25, 2025Updated 9 months ago
Alternatives and similar repositories for awesome-captioning-evaluation
Users that are interested in awesome-captioning-evaluation are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Recurrence Meets Transformers for Universal Multimodal Retrieval☆15Dec 15, 2025Updated 8 months ago
- [ICCV 2025] What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models☆16Nov 3, 2025Updated 9 months ago
- This is the official repository for the paper "Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction". ICCV …☆27May 13, 2026Updated 3 months ago
- [ICCV 2023] With a Little Help from your own Past: Prototypical Memory Networks for Image Captioning.☆19Jun 7, 2024Updated 2 years ago
- 我在校园的各项API,自动运行脚本,支持多人☆12Jun 28, 2022Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official implementation of the Composed Image Retrieval using Pretrained LANguage Transformers (CIRPLANT) | ICCV 2021 - Image Retrieval o…☆39Jun 26, 2024Updated 2 years ago
- Implementing NVLabs C-RADIOv3 Embeddings Model as Remotely Sourced Zoo Model for FiftyOne☆27Feb 3, 2026Updated 6 months ago
- large scale pre-training VLMs☆25Jul 6, 2026Updated last month
- [CVPR'25] AVF-MAE++ : Scaling Affective Video Facial Masked Autoencoders via Efficient Audio-Visual Self-Supervised Learning☆23Jun 11, 2026Updated 2 months ago
- Generating Soil Moisture (SM) or Evapotranspiration (ET) maps from satellite images☆10Aug 23, 2019Updated 7 years ago
- Microsoft question-answering dataset☆10Jun 16, 2023Updated 3 years ago
- The Easiest Way to Run Commands as Systemd Services☆11Updated this week
- [ACM MM 2025] This repository is the official implementation of the paper "Motion Matters: Motion-guided Modulation Network for Skeleton-…☆22Nov 28, 2025Updated 9 months ago
- The Land-Diffuser is a novel application of the Denoising Diffusion Probabilistic Model (DDPM) in the realm of 3D Talking Head generation…☆13Dec 23, 2023Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ICLR 2025] - Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion☆70Nov 30, 2025Updated 9 months ago
- The code for the Ensemble everything everywhere: Multi-scale aggregation for adversarial robustness paper☆23Nov 8, 2024Updated last year
- [CVPR24 Highlights] Polos: Multimodal Metric Learning from Human Feedback for Image Captioning☆33Jun 12, 2026Updated 2 months ago
- [CVPR 2025 Highlight] Official implementation of HySAC, a hyperbolic safety-aware vision-language model for safer multimodal retrieval an…☆31Apr 8, 2025Updated last year
- [CVPR '26] CaptionQA: Is Your Caption as Useful as the Image Itself?☆38Mar 3, 2026Updated 5 months ago
- PyTorch code for the paper: "Perceive, Transform, and Act: Multi-Modal Attention Networks for Vision-and-Language Navigation"☆19Aug 5, 2021Updated 5 years ago
- Official repo of the OpenApePose dataset☆13Feb 6, 2024Updated 2 years ago
- An Arena-style Automated Evaluation Benchmark for Detailed Captioning☆59Jun 1, 2025Updated last year
- [ICLR 2025] Causal Graphical Models for Vision-Language Compositional Understanding☆10Apr 15, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [NO LONGER MAINTAINED, SUPERSEDED BY https://github.com/trueagi-io/pln-experimental and https://github.com/trueagi-io/PLN]. Probabilisti…☆16Sep 20, 2025Updated 11 months ago
- FeelingBlue: A Corpus for Understanding the Emotional Connotation of Color in Context, accepted at TACL 2022, presented at ACL 2023☆13Dec 28, 2023Updated 2 years ago
- ☆17Aug 23, 2025Updated last year
- SOMOSPIE (Soil Moisture Spatial Inference Engine) consists of a Jupyter Notebook and a suite of machine learning methods to process input…☆16Updated this week
- Core implementation of gl-shader without parser dependencies☆22Apr 24, 2016Updated 10 years ago
- A currency rate converter App.☆15Sep 5, 2019Updated 6 years ago
- This repository contains the dataset and code for our ACL'23 publication: "MatSci-NLP: Evaluating Scientific Language Models on Materials…☆17Nov 21, 2023Updated 2 years ago
- [ICCV 2025] - Image Intrinsic Scale Assessment: Bridging the Gap Between Quality and Resolution☆17Aug 16, 2025Updated last year
- Understanding Rare Spurious Correlations in Neural Network☆12Jun 5, 2022Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆12Apr 3, 2026Updated 4 months ago
- TUI application for viewing the status of GPU allocations on a Slurm cluster☆11Dec 11, 2023Updated 2 years ago
- Official code of "Where are my Neighbors? Exploiting Patches Relations in Self-Supervised Vision Transformer", Guglielmo Camporese, Elena…☆21Dec 14, 2022Updated 3 years ago
- Mitigating Open-Vocabulary Caption Hallucinations (EMNLP 2024)☆19Oct 18, 2024Updated last year
- Training A Small Emotional Vision Language Model for Visual Art Comprehension☆17Jul 26, 2024Updated 2 years ago
- Farm Scale Soil Moisture using Remote Sensing and Water Budget Products☆23Mar 30, 2024Updated 2 years ago
- This is the code for our ACL 2021 paper entitled eMLM: A New Pre-training Objective for Emotion Related Tasks☆15Sep 7, 2022Updated 3 years ago