VLM2-Bench [ACL 2025 Main]: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues
☆45May 20, 2025Updated last year
Alternatives and similar repositories for VLM2-Bench
Users that are interested in VLM2-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The official codebase for our paper "NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems"☆24Aug 26, 2026Updated 3 weeks ago
- Official Repository for our paper: PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems☆40Updated this week
- ☆49Aug 12, 2026Updated last month
- This is the official implementation for MA-LoT.☆20Aug 4, 2025Updated last year
- Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to …☆74Jan 28, 2026Updated 7 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆12May 15, 2025Updated last year
- ☆25Aug 2, 2024Updated 2 years ago
- Public code repo for EMNLP 2024 Findings paper "MACAROON: Training Vision-Language Models To Be Your Engaged Partners"☆14Sep 28, 2024Updated last year
- RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment☆18Dec 19, 2024Updated last year
- A toolkit for systematically understanding the concepts encoded in Sparse Autoencoders.☆19Apr 5, 2026Updated 5 months ago
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis☆40Sep 15, 2026Updated last week
- Code for paper "Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models."☆55Oct 19, 2024Updated last year
- Official Code of UniFine☆11Jan 17, 2024Updated 2 years ago
- [NeurIPS 2024 D&B] Evaluating Copyright Takedown Methods for Language Models☆17Jul 17, 2024Updated 2 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Official Repository of Personalized Visual Instruct Tuning☆34Mar 6, 2025Updated last year
- Repository of <FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models> in EMNLP 2026 Findings☆76Aug 24, 2026Updated 3 weeks ago
- CyberGym-E2E is a large-scale benchmark built from real-world vulnerabilities in widely used open-source projects to evaluate AI agents' …☆75Sep 3, 2026Updated 2 weeks ago
- This repo contains script to download MUSIC dataset from youtube☆13Jan 19, 2024Updated 2 years ago
- Code and data for the paper: IntentionQA: A Benchmark for Evaluating Purchase Intention Comprehension Abilities of Large Language Models …☆12Apr 27, 2024Updated 2 years ago
- [ACL 2024] "Understanding and Patching Compositional Reasoning in LLMs"☆14Aug 8, 2026Updated last month
- ☆18Jan 6, 2025Updated last year
- [ICML 2024] Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.☆91Jan 19, 2025Updated last year
- The released data for paper "Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models".☆34Sep 16, 2023Updated 3 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Image Textualization: An Automatic Framework for Generating Rich and Detailed Image Descriptions (NeurIPS 2024)☆173Jul 30, 2024Updated 2 years ago
- ☆19Oct 14, 2024Updated last year
- Official Implementation of VoxTracer (MM' 23)☆12Oct 27, 2023Updated 2 years ago
- Official implementation of CVPR 2024 paper "Prompt Learning via Meta-Regularization".☆31Mar 10, 2025Updated last year
- Train vector quantized CLIP models using pytorch lightning☆21Jul 14, 2024Updated 2 years ago
- [ICLR 2026] Dancing in Chains: Strategic Persuasion in Academic Rebuttal via Theory of Mind (If you find this useful, please give us a st…☆57Apr 28, 2026Updated 4 months ago
- Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios☆29Sep 30, 2025Updated 11 months ago
- 18级武汉大学国家网络安全学院暑期实训备份☆11Jul 18, 2019Updated 7 years ago
- ☆13Jan 16, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Implementation of "A Large-Scale Study of Probabilistic Calibration in Neural Network Regression" (ICML 2023)☆11Oct 7, 2025Updated 11 months ago
- [ACL'24 Findings] Official code for "TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback"☆12Dec 6, 2024Updated last year
- ☆18Jul 6, 2026Updated 2 months ago
- ☆12Jun 12, 2024Updated 2 years ago
- ☆30Jul 8, 2026Updated 2 months ago
- Official code and dataset for our NAACL 2024 paper: DialogCC: An Automated Pipeline for Creating High-Quality Multi-modal Dialogue Datase…☆13Jun 24, 2024Updated 2 years ago
- Official repository of Generating Multiple-Length Summaries via Reinforcement Learning for Unsupervised Sentence Summarization [EMNLP'22 …☆10May 20, 2023Updated 3 years ago