VLM2-Bench [ACL 2025 Main]: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues
☆45May 20, 2025Updated last year
Alternatives and similar repositories for VLM2-Bench
Users that are interested in VLM2-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The official code repository for the paper "CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments…☆35Jun 14, 2026Updated 2 months ago
- The official codebase for our paper "NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems"☆24Aug 26, 2026Updated last week
- Official Repository for our paper: PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems☆40Jul 16, 2026Updated last month
- This is the official implementation for MA-LoT.☆20Aug 4, 2025Updated last year
- Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to …☆74Jan 28, 2026Updated 7 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆12May 15, 2025Updated last year
- ☆25Aug 2, 2024Updated 2 years ago
- One implementation of the paper "Controllable Neural Dialogue Summarization with Personal Named Entity Planning" (EMNLP 2022).☆18Nov 9, 2023Updated 2 years ago
- Public code repo for EMNLP 2024 Findings paper "MACAROON: Training Vision-Language Models To Be Your Engaged Partners"☆14Sep 28, 2024Updated last year
- [ICML 2026] XSkill: Continual Learning from Experience and Skills in Multimodal Agents☆261May 13, 2026Updated 3 months ago
- Benchmarking multimodal agents on realistic, ultra-challenging visual scenarios requiring long-horizon hybrid tool use.☆73Mar 10, 2026Updated 5 months ago
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis☆38Jul 10, 2026Updated last month
- Model calibration in CLIP Adapters☆20Aug 19, 2024Updated 2 years ago
- Code for paper "Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models."☆55Oct 19, 2024Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Repository of <FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models> in EMNLP 2026 Findings☆75Aug 24, 2026Updated last week
- Official Repository of Personalized Visual Instruct Tuning☆34Mar 6, 2025Updated last year
- Code and data for the paper: IntentionQA: A Benchmark for Evaluating Purchase Intention Comprehension Abilities of Large Language Models …☆12Apr 27, 2024Updated 2 years ago
- [ACL 2024] "Understanding and Patching Compositional Reasoning in LLMs"☆14Aug 8, 2026Updated 3 weeks ago
- ☆18Jan 6, 2025Updated last year
- [ICML 2024] Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.☆91Jan 19, 2025Updated last year
- The opensoure repository of FuzzLLM☆37May 4, 2024Updated 2 years ago
- The released data for paper "Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models".☆34Sep 16, 2023Updated 2 years ago
- Image Textualization: An Automatic Framework for Generating Rich and Detailed Image Descriptions (NeurIPS 2024)☆173Jul 30, 2024Updated 2 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆19Oct 14, 2024Updated last year
- AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation (EMNLP 2024 Findings)☆19Dec 30, 2024Updated last year
- Official Implementation of VoxTracer (MM' 23)☆12Oct 27, 2023Updated 2 years ago
- Train vector quantized CLIP models using pytorch lightning☆21Jul 14, 2024Updated 2 years ago
- A Good Neighbor, A Found Treasure: Mining Treasured Neighbors for Knowledge Graph Entity Typing. EMNLP 2022☆11Feb 1, 2023Updated 3 years ago
- [ICLR 2026] Dancing in Chains: Strategic Persuasion in Academic Rebuttal via Theory of Mind (If you find this useful, please give us a st…☆57Apr 28, 2026Updated 4 months ago
- Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios☆28Sep 30, 2025Updated 11 months ago
- Implementation of "A Large-Scale Study of Probabilistic Calibration in Neural Network Regression" (ICML 2023)☆11Oct 7, 2025Updated 10 months ago
- The re-implementation of <End-to-End Lane Marker Detection via Row-wise Classification>☆15Sep 21, 2020Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ACL'24 Findings] Official code for "TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback"☆12Dec 6, 2024Updated last year
- This is the code for "TARGET: Federated Class-Continual Learning via Exemplar-Free Distillation" (ICCV 2023)☆52Apr 30, 2024Updated 2 years ago
- ☆12Jun 12, 2024Updated 2 years ago
- Official code and dataset for our NAACL 2024 paper: DialogCC: An Automated Pipeline for Creating High-Quality Multi-modal Dialogue Datase…☆13Jun 24, 2024Updated 2 years ago
- Code and data for the ACL 2024 Findings paper "Do LVLMs Understand Charts? Analyzing and Correcting Factual Errors in Chart Captioning"☆27Jun 5, 2024Updated 2 years ago
- Source code for paper "PRiSM: Enhancing Low-Resource Document-Level Relation Extraction with Relation-Aware Score Calibration", Findings …☆11Jun 20, 2025Updated last year
- Official implementation of ICLR 2026: Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement☆15May 24, 2026Updated 3 months ago