[CVPR '26] CaptionQA: Is Your Caption as Useful as the Image Itself?
☆38Mar 3, 2026Updated 5 months ago
Alternatives and similar repositories for CaptionQA
Users that are interested in CaptionQA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16Jun 23, 2026Updated last month
- HallE-Control: Controlling Object Hallucination in LMMs☆32Apr 10, 2024Updated 2 years ago
- [ICML 2026] Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions☆53Jun 29, 2026Updated last month
- [ICLR 2026] An official implementation of "CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning"☆227Jun 23, 2026Updated last month
- ☆12Jul 4, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Towards Memorization-Free Diffusion Models (CVPR2024) Codebase☆11Jun 2, 2024Updated 2 years ago
- [ICCV 2023] With a Little Help from your own Past: Prototypical Memory Networks for Image Captioning.☆19Jun 7, 2024Updated 2 years ago
- [ECCV2024] Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models☆20Jul 17, 2024Updated 2 years ago
- ☆14Apr 1, 2023Updated 3 years ago
- [ECCV 2024] The first zero-shot setting for spatio-temporal video grounding.☆11Jul 16, 2024Updated 2 years ago
- ☆12Sep 30, 2024Updated last year
- ☆14Dec 25, 2024Updated last year
- ASID-Caption: Attribute-Structured and Quality-Verified Audiovisual Instruction Dataset and Training Pipeline for Fine-Grained Video Unde…☆68Mar 3, 2026Updated 5 months ago
- Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations☆23Jul 12, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning☆17May 27, 2026Updated 2 months ago
- [ICCV'23] PAINet: Parallel Attention Interaction Network for Few-shot Skeleton-based Action Recognition☆11Oct 14, 2023Updated 2 years ago
- ☆32Jul 29, 2024Updated 2 years ago
- Research code for ACL2024 paper: "Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline"☆42Dec 27, 2024Updated last year
- [ICLR'26] Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs☆100Jan 26, 2026Updated 6 months ago
- ☆13Apr 30, 2025Updated last year
- [NeurIPS 2025 Spotlight] Official PyTorch implementation of Vgent☆50Nov 30, 2025Updated 8 months ago
- [ECCV 2026] Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?☆97Jul 13, 2025Updated last year
- [ACL 2025] Official code for ''Learning to Reason from Feedback at Test-Time''.☆13May 16, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Math-VR Benchmark & CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images☆63Nov 4, 2025Updated 9 months ago
- Open Source Face Recognition Performance Evaluation Package☆20Oct 7, 2019Updated 6 years ago
- A cross-chatting-platform bot framework.一个跨聊天平台的支持多语言开发的跨平台机器人框架。☆16Jan 25, 2023Updated 3 years ago
- ☆20May 3, 2025Updated last year
- Developing adversarial examples and showing their semantic generalization for the OpenAI CLIP model (https://github.com/openai/CLIP)☆25Mar 6, 2021Updated 5 years ago
- The official implementation of NOSA☆20Jun 11, 2026Updated 2 months ago
- [ICCV2025] The official code of "DreamRelation: Relation-Centric Video Customization"☆27Feb 4, 2026Updated 6 months ago
- Spatial Temporal Graph Convolutional Networks (ST-GCN) for Skeleton-Based Action Recognition in PyTorch☆18Jan 25, 2018Updated 8 years ago
- [ECCV 2026] Towards Scalable Pre-training of Visual Tokenizers for Generation☆500Apr 15, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆189Jun 27, 2025Updated last year
- ☆23May 3, 2025Updated last year
- ☆18Mar 1, 2026Updated 5 months ago
- Official implementation of paper ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding☆40Mar 16, 2025Updated last year
- 👋This repo is used to host my personal blog, check https://martinlwx.github.io☆15Jul 8, 2026Updated last month
- ☆22Jan 26, 2024Updated 2 years ago
- [ICML 2024] CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers☆34Dec 30, 2024Updated last year