☆24Jun 18, 2025Updated last year
Alternatives and similar repositories for ViCrit
Users that are interested in ViCrit are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆15Mar 30, 2025Updated last year
- ☆105Jun 10, 2025Updated last year
- VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection☆27May 31, 2025Updated last year
- ☆32Feb 8, 2024Updated 2 years ago
- ☆11Jul 31, 2022Updated 4 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Official Repository: A Comprehensive Benchmark for Logical Reasoning in MLLMs☆45Jun 17, 2025Updated last year
- [ECCV 2026] Official implementation of "TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning"☆25Feb 8, 2026Updated 7 months ago
- Official implementation of Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal Learning☆28Oct 30, 2024Updated last year
- [ACL 2023] VSTAR is a multimodal dialogue dataset with scene and topic transition information☆16Oct 27, 2024Updated last year
- [CVPR'24] HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(…☆343Oct 14, 2025Updated 11 months ago
- ☆100Mar 29, 2019Updated 7 years ago
- Official repository for the A-OKVQA dataset☆118May 8, 2024Updated 2 years ago
- ☆15Nov 13, 2025Updated 10 months ago
- Developer project for getting basic API integrations working in under 5 minutes☆11May 22, 2026Updated 3 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [ACL-main-2026]We introduce Chart2Code, the first user-driven, hierarchical benchmark that systematically evaluates Large Multimodal Mode…☆31Aug 4, 2026Updated last month
- [ICCV 2025] MMReason, MLLMs, step by step, reasoning benchmark, AGI☆15Apr 25, 2026Updated 4 months ago
- [NeurIPS 2025] NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation☆112Sep 18, 2025Updated last year
- Code for the WACV 2020 paper "Answering Questions about Data Visualizations using Efficient Bimodal Fusion"☆14Jun 22, 2021Updated 5 years ago
- [AAAI 2024] VSFormer: Visual-Spatial Fusion Transformer for Correspondence Pruning☆16Apr 7, 2024Updated 2 years ago
- ☆86May 2, 2026Updated 4 months ago
- VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation☆20Jun 2, 2025Updated last year
- [ICLR 2026]🚀ReVisual-R1 is a 7B open-source multimodal language model that follows a three-stage curriculum—cold-start pre-training, mul…☆210Dec 10, 2025Updated 9 months ago
- A curated list of all awesome pygames created by Agneay B Nair☆12Apr 28, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- CVPR '26 Highlight☆24May 6, 2026Updated 4 months ago
- An LLM-free Multi-dimensional Benchmark for Multi-modal Hallucination Evaluation☆173Jan 15, 2024Updated 2 years ago
- [ICCV 2023] HiLo: Exploiting High Low Frequency Relations for Unbiased Panoptic Scene Graph Generation☆38Jan 25, 2024Updated 2 years ago
- Official Implementation for "ESCAPE: Encoding Super-keypoints for Category-Agnostic Pose Estimation", CVPR 2024.☆10Jun 17, 2024Updated 2 years ago
- [CVPR 2025] Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention☆69Jul 16, 2024Updated 2 years ago
- ☆12Sep 19, 2021Updated 5 years ago
- ☆22Jan 26, 2025Updated last year
- Code and data release for the paper "Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Align…☆19Apr 5, 2024Updated 2 years ago
- Computer-Use Agents as Judges for Generative UI☆45Nov 27, 2025Updated 9 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code for "AVG-LLaVA: A Multimodal Large Model with Adaptive Visual Granularity"☆33Oct 12, 2024Updated last year
- AutoHallusion Codebase (EMNLP 2024)☆22Dec 6, 2024Updated last year
- [ACL 2025] Analyzing LLMs' Multilingual Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations☆19Oct 18, 2025Updated 11 months ago
- A unified approach to explain conditional text generation models. Pytorch. The code of paper "Local Explanation of Dialogue Response Gene…☆16Mar 21, 2022Updated 4 years ago
- ☆31May 21, 2026Updated 3 months ago
- Repository for the paper: Teaching Structured Vision & Language Concepts to Vision & Language Models☆47Sep 25, 2023Updated 2 years ago
- ☆81Oct 27, 2023Updated 2 years ago