VLM2-Bench [ACL 2025 Main]: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues
☆45May 20, 2025Updated last year
Alternatives and similar repositories for VLM2-Bench
Users that are interested in VLM2-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The official code repository for the paper "CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments…☆33Jun 14, 2026Updated last month
- Official Repository for our paper: PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems☆38Jul 16, 2026Updated last week
- ☆43Dec 19, 2025Updated 7 months ago
- This is the official implementation for MA-LoT.☆20Aug 4, 2025Updated 11 months ago
- Benchmark for Hetergeneous Federated Learning by MARS Group at the Wuhan University, led by Prof. Mang Ye.☆19May 29, 2023Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to …☆72Jan 28, 2026Updated 5 months ago
- ☆12May 15, 2025Updated last year
- ☆25Aug 2, 2024Updated last year
- One implementation of the paper "Controllable Neural Dialogue Summarization with Personal Named Entity Planning" (EMNLP 2022).☆18Nov 9, 2023Updated 2 years ago
- Public code repo for EMNLP 2024 Findings paper "MACAROON: Training Vision-Language Models To Be Your Engaged Partners"☆14Sep 28, 2024Updated last year
- [ICML 2026] XSkill: Continual Learning from Experience and Skills in Multimodal Agents☆239May 13, 2026Updated 2 months ago
- RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment☆18Dec 19, 2024Updated last year
- A toolkit for systematically understanding the concepts encoded in Sparse Autoencoders.☆20Apr 5, 2026Updated 3 months ago
- Benchmarking multimodal agents on realistic, ultra-challenging visual scenarios requiring long-horizon hybrid tool use.☆67Mar 10, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis☆33Jul 10, 2026Updated 2 weeks ago
- Model calibration in CLIP Adapters☆20Aug 19, 2024Updated last year
- Code for paper "Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models."☆54Oct 19, 2024Updated last year
- [NeurIPS 2024 D&B] Evaluating Copyright Takedown Methods for Language Models☆17Jul 17, 2024Updated 2 years ago
- Repository of <FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models>☆75Jan 8, 2026Updated 6 months ago
- Official Repository of Personalized Visual Instruct Tuning☆34Mar 6, 2025Updated last year
- CyberGym-E2E is a large-scale benchmark built from real-world vulnerabilities in widely used open-source projects to evaluate AI agents' …☆29Jun 25, 2026Updated 3 weeks ago
- This repo contains script to download MUSIC dataset from youtube☆12Jan 19, 2024Updated 2 years ago
- Code and data for the paper: IntentionQA: A Benchmark for Evaluating Purchase Intention Comprehension Abilities of Large Language Models …☆12Apr 27, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ACL 2024] "Understanding and Patching Compositional Reasoning in LLMs"☆14Aug 28, 2024Updated last year
- [ICML 2024] Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.☆90Jan 19, 2025Updated last year
- The released data for paper "Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models".☆34Sep 16, 2023Updated 2 years ago
- Image Textualization: An Automatic Framework for Generating Rich and Detailed Image Descriptions (NeurIPS 2024)☆172Jul 30, 2024Updated last year
- ☆19Oct 14, 2024Updated last year
- AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation (EMNLP 2024 Findings)☆18Dec 30, 2024Updated last year
- Official Implementation of VoxTracer (MM' 23)☆12Oct 27, 2023Updated 2 years ago
- Official implementation of CVPR 2024 paper "Prompt Learning via Meta-Regularization".☆31Mar 10, 2025Updated last year
- A Good Neighbor, A Found Treasure: Mining Treasured Neighbors for Knowledge Graph Entity Typing. EMNLP 2022☆11Feb 1, 2023Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios☆27Sep 30, 2025Updated 9 months ago
- LatentMAS with kNN kv cache pruning | up to 40% more memory efficient and 30% faster☆19Dec 10, 2025Updated 7 months ago
- ☆13Jan 16, 2025Updated last year
- Implementation of "A Large-Scale Study of Probabilistic Calibration in Neural Network Regression" (ICML 2023)☆11Oct 7, 2025Updated 9 months ago
- The re-implementation of <End-to-End Lane Marker Detection via Row-wise Classification>☆14Sep 21, 2020Updated 5 years ago
- [ACL'24 Findings] Official code for "TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback"☆12Dec 6, 2024Updated last year
- ☆12Jun 12, 2024Updated 2 years ago