[CVPR '26] CaptionQA: Is Your Caption as Useful as the Image Itself?
☆39Mar 3, 2026Updated 6 months ago
Alternatives and similar repositories for CaptionQA
Users that are interested in CaptionQA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16Jun 23, 2026Updated 2 months ago
- HallE-Control: Controlling Object Hallucination in LMMs☆32Apr 10, 2024Updated 2 years ago
- [ICML 2026] Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions☆56Jun 29, 2026Updated 2 months ago
- [ICCV 2023] With a Little Help from your own Past: Prototypical Memory Networks for Image Captioning.☆19Jun 7, 2024Updated 2 years ago
- [NeurIPS 2024] PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications☆22Nov 4, 2024Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ECCV2024] Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models☆20Jul 17, 2024Updated 2 years ago
- [COLM'25] Official implementation of the Law of Vision Representation in MLLMs☆179Oct 6, 2025Updated 11 months ago
- ☆14Apr 1, 2023Updated 3 years ago
- ☆12Dec 6, 2024Updated last year
- Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries☆47Nov 19, 2025Updated 9 months ago
- [ECCV 2024] The first zero-shot setting for spatio-temporal video grounding.☆11Jul 16, 2024Updated 2 years ago
- ☆12Sep 30, 2024Updated last year
- Task-adaptive Spatial-Temporal Video Sampler for Few-shot Action Recognition☆14Dec 22, 2022Updated 3 years ago
- ASID-Caption: Attribute-Structured and Quality-Verified Audiovisual Instruction Dataset and Training Pipeline for Fine-Grained Video Unde…☆70Mar 3, 2026Updated 6 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations☆23Jul 12, 2026Updated last month
- DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning☆17May 27, 2026Updated 3 months ago
- [ICCV'23] PAINet: Parallel Attention Interaction Network for Few-shot Skeleton-based Action Recognition☆11Oct 14, 2023Updated 2 years ago
- ☆32Jul 29, 2024Updated 2 years ago
- Research code for ACL2024 paper: "Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline"☆42Dec 27, 2024Updated last year
- [ICLR'26] Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs☆101Jan 26, 2026Updated 7 months ago
- [NeurIPS 2025 Spotlight] Official PyTorch implementation of Vgent☆50Nov 30, 2025Updated 9 months ago
- [ECCV 2026] Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?☆99Jul 13, 2025Updated last year
- ☆11Mar 16, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ICLR 2026 Oral] Official Implementation of the paper "MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interactio…☆22Jul 2, 2026Updated 2 months ago
- [ACL 2025] Official code for ''Learning to Reason from Feedback at Test-Time''.☆13May 16, 2025Updated last year
- Math-VR Benchmark & CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images☆64Nov 4, 2025Updated 10 months ago
- [CVPR 2024] This repository includes the official implementation our paper "Revisiting Adversarial Training at Scale"☆20Apr 21, 2024Updated 2 years ago
- ☆20May 3, 2025Updated last year
- The official implementation of NOSA☆20Jun 11, 2026Updated 2 months ago
- [ICCV2025] The official code of "DreamRelation: Relation-Centric Video Customization"☆26Feb 4, 2026Updated 7 months ago
- We introduce CausalVQA, a benchmark dataset for video question answering (VQA) composed of question-answer pairs that probe models’ under…☆62Aug 18, 2025Updated last year
- [ECCV 2026] Towards Scalable Pre-training of Visual Tokenizers for Generation☆507Apr 15, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆190Jun 27, 2025Updated last year
- ☆23May 3, 2025Updated last year
- [ACL 2025] "World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning." https://arxiv.org/abs/2503.1…☆18Jul 22, 2025Updated last year
- ☆18Mar 1, 2026Updated 6 months ago
- Official implementation of paper ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding☆40Mar 16, 2025Updated last year
- [NeurIPS 2025] Bag of Tricks for Inference-time Computation of LLM Reasoning☆16Sep 20, 2025Updated 11 months ago
- ☆30Aug 29, 2026Updated last week