[ACL 2026] WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering.
☆15Jul 25, 2026Updated 3 weeks ago
Alternatives and similar repositories for WikiSeeker
Users that are interested in WikiSeeker are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACM-MM 2025 Workshop] More Is Better: A MoE-Based Emotion Recognition Framework with Human Preference Alignment.☆25Nov 25, 2025Updated 8 months ago
- [Neurocomputing] Efficient Redundancy Reduction for Open-Vocabulary Semantic Segmentation☆26Dec 21, 2025Updated 8 months ago
- [ICML 2026] Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions☆25Feb 11, 2026Updated 6 months ago
- ☆19Mar 9, 2026Updated 5 months ago
- Python基于CRNN&CTPN的文本检测系统(源码&教程)☆10Nov 14, 2023Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- 提供一种基于深度学习的光纤传感水声信号识别方法,该方法降低了光纤传感水声信号识别的难度,通过最优聚类模型,将无监督学习方式转化为有监督学习的方式,使识别未知的目标事件信号成为可能;以光纤传感系统自身固有噪声信号分解分量作为训练数据,构建开集识别网络,可用于识别任意不属于系统…☆12Aug 30, 2022Updated 3 years ago
- Official data and code for the paper "VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents".☆15Mar 18, 2026Updated 5 months ago
- ☆12Sep 27, 2024Updated last year
- [ICML 2025 Spotlight] Official implementation of the paper: Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Mod…☆19Sep 1, 2025Updated 11 months ago
- [CVPR 2025] Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering