π₯ [ICLR 2025] Official Benchmark Toolkits for "Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark"
β44Nov 21, 2025Updated 10 months ago
Alternatives and similar repositories for vhs_benchmark
Users that are interested in vhs_benchmark are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π₯ [ICLR 2025] Official PyTorch Model "Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark"β26Feb 9, 2025Updated last year
- We introduce new approach, Token Reduction using CLIP Metric (TRIM), aimed at improving the efficiency of MLLMs without sacrificing theirβ¦β23Jan 11, 2026Updated 8 months ago
- π₯ [ICML 2026] Official implementation of "Are LRMs Interruptible?"β20Jun 18, 2026Updated 3 months ago
- A Comprehensive Benchmark for Robust Multi-image Understandingβ21Sep 4, 2024Updated 2 years ago
- Echo: "Constantly Improving Image Models Need Constantly Improving Benchmarks" (ICLR 2026)β20Jan 29, 2026Updated 7 months ago
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- π Official pytorch implementation of "D2ADA: Dynamic Density-aware Active Domain Adaptation for Semantic Segmentation. Wu et al. ECCV 20β¦β25Feb 2, 2023Updated 3 years ago
- VELMA agent for VLN in Street Viewβ31Sep 29, 2023Updated 2 years ago
- π₯ [NeurIPS 2025] Official implementation of "Generate, but Verify: Reducing Visual Hallucination in Vision-Language Models with Retrospeβ¦β59Jan 22, 2026Updated 8 months ago
- β22Apr 24, 2025Updated last year
- The goal of this experiment is to take articles and certain metadata and group them by topic.β11Apr 14, 2016Updated 10 years ago
- Ask&Confirm: Active Detail Enriching for Cross-Modal Retrieval with Partial Query (ICCV2021)β20Dec 4, 2021Updated 4 years ago
- VideoNIAH: A Flexible Synthetic Method for Benchmarking Video MLLMsβ57Mar 9, 2025Updated last year
- Official code repository for the main conference paper in EMNLP 2022: SubeventWriter: Iterative Sub-event Sequence Generation with Cohereβ¦β11Oct 16, 2022Updated 3 years ago
- Official Repository of VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agentsβ115May 3, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Implementation of the model: "Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models" in PyTorchβ29Aug 29, 2026Updated 3 weeks ago
- Official code for CVPR 2024 paper, "Audio-Visual Segmentation via Unlabeled Frame Exploitation""β19Jul 7, 2024Updated 2 years ago
- [ACL 2025 Long Main] Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinionsβ48Apr 21, 2025Updated last year
- CLAIR: A (surprisingly) simple semantic text metric with large language models.β23Jan 28, 2024Updated 2 years ago
- π [NeurIPS24] Make Vision Matter in Visual-Question-Answering (VQA)! Introducing NaturalBench, a vision-centric VQA benchmark (NeurIPS'2β¦β90Jun 24, 2025Updated last year
- Codebase, data and models for the Headline Grouping paper at NAACL2021β12Oct 2, 2022Updated 3 years ago
- [COLM'25] Official implementation of the Law of Vision Representation in MLLMsβ179Oct 6, 2025Updated 11 months ago
- Code and Data for ACL 2023 paper I Spy a Metaphor: Large Language Models and Diffusion Models Co-Create Visual Metaphorsβ17Jun 7, 2023Updated 3 years ago
- [NeurIPS 2024] Needle In A Multimodal Haystack (MM-NIAH): A comprehensive benchmark designed to systematically evaluate the capability ofβ¦β130Nov 25, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Holistic evaluation of multimodal foundation modelsβ48Aug 11, 2024Updated 2 years ago
- [ACM Computing Survey 2025] Vertical Federated Learning for Effectiveness, Security, Applicability: A Survey, by MARS Group at Wuhan Univβ¦β23Apr 9, 2025Updated last year
- [EMNLP 2024] mDPO: Conditional Preference Optimization for Multimodal Large Language Models.β91Nov 10, 2024Updated last year
- Code release for "MORE: Multi-mOdal REtrieval Augmented Generative Commonsense Reasoning"β11Oct 11, 2024Updated last year
- Implementation of the Unsupervised learning by predicting noise paperβ14Jan 18, 2018Updated 8 years ago
- learning materials of driveos from nvidia drive sdk.β13Jun 10, 2026Updated 3 months ago
- An LLM-free Multi-dimensional Benchmark for Multi-modal Hallucination Evaluationβ174Jan 15, 2024Updated 2 years ago
- β28Jul 22, 2026Updated 2 months ago
- Repository for paper Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queriesβ13Jul 16, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- a set of tools for computer vision processingβ18Jul 9, 2016Updated 10 years ago
- Official code repository for Findings of EMNLP 2022 paper: PseudoReasoner: Leveraging Pseudo Labels for Commonsense Knowledge Base Populaβ¦β11Oct 18, 2022Updated 3 years ago
- Iterate on LLM-based structured generation forward and backwardβ24Mar 20, 2025Updated last year
- [CVPR23 Highlight] CREPE: Can Vision-Language Foundation Models Reason Compositionally?β36Apr 27, 2023Updated 3 years ago
- Code release for "UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity"β86Sep 15, 2026Updated last week
- β117Sep 13, 2025Updated last year
- β34May 14, 2025Updated last year