v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
☆21Sep 16, 2026Updated 3 weeks ago
Alternatives and similar repositories for v1
Users that are interested in v1 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2025] This is the official repository for VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Se…☆15Oct 3, 2026Updated last week
- ☆18Nov 8, 2025Updated 11 months ago
- Corpus to accompany: "Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding"☆11Apr 11, 2025Updated last year
- ☆11Updated this week
- ☆14Jun 20, 2023Updated 3 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Wikidata properties☆10Aug 12, 2026Updated last month
- ☆30Mar 13, 2024Updated 2 years ago
- This is the implementation of CounterCurate, the data curation pipeline of both physical and semantic counterfactual image-caption pairs.☆19Jun 27, 2024Updated 2 years ago
- Code for the arxiv paper: Complex Claim Verification with Evidence Retrieved in the Wild☆15Nov 27, 2023Updated 2 years ago
- [ICCV 2025] Auto Interpretation Pipeline and many other functionalities for Multimodal SAE Analysis.☆202Sep 26, 2025Updated last year
- Download Web-10K data by querying Bing Image Search☆10Feb 1, 2022Updated 4 years ago
- Official pytorch implementation of "Towards Practical Plug-and-Play Diffusion Models" in CVPR2023☆22Jul 22, 2023Updated 3 years ago
- Ultra-minimal autoregressive diffusion model for image generation☆21Jul 20, 2026Updated 2 months ago
- ChartSum is a large scale benchmark for automatic chart to text summarization☆11Jul 20, 2023Updated 3 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- [CVPR 2026] Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens☆300Aug 2, 2025Updated last year
- Pytorch Code for "Unified Coarse-to-Fine Alignment for Video-Text Retrieval" (ICCV 2023)☆66Jun 7, 2024Updated 2 years ago
- ☆11Dec 9, 2025Updated 10 months ago
- VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning☆38May 9, 2026Updated 5 months ago
- Train vector quantized CLIP models using pytorch lightning☆21Jul 14, 2024Updated 2 years ago
- ☆105Jun 10, 2025Updated last year
- ☆12Jun 18, 2024Updated 2 years ago
- [AAAI'25]: Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP☆23Aug 5, 2025Updated last year
- Code for Dayal Kalra's research internship on scalable curvature measures for neural networks.☆32Feb 3, 2026Updated 8 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ShopEasy is a containerized e‑commerce boilerplate with an Angular SPA and Spring Boot API (JWT‑secure), backed by MySQL and Kafka, deplo…☆15Sep 5, 2025Updated last year
- Repo for the paper: Towards Few-shot Entity Recognition in Document Images:A Label-aware Sequence-to-Sequence Framework☆14May 31, 2023Updated 3 years ago
- HOCR Specification Python Parser☆12Sep 23, 2015Updated 11 years ago
- ☆13Aug 14, 2022Updated 4 years ago
- Official code for NeurIPS 2025 paper "GRIT: Teaching MLLMs to Think with Images"☆192Jan 16, 2026Updated 8 months ago
- ☆27Feb 9, 2023Updated 3 years ago
- ☆78Apr 9, 2026Updated 6 months ago
- [NeurIPS 2025] MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning☆108Sep 19, 2025Updated last year
- A general purpose gpu computation library in odin☆12May 9, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Code for EMNLP2021 paper “Transductive Learning for Unsupervised Text Style Transfer”☆12Sep 19, 2021Updated 5 years ago
- This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception o…☆29Jul 9, 2025Updated last year
- Code for paper 'Zero-Shot Scene Graph Generation via Triplet Calibration and Reduction' (TOMM 2023)☆10Sep 6, 2025Updated last year
- In this codebase we establish a benchmark for egocentric user adaptation based on Ego4d.First, we start from a population model which ha…☆14Jul 24, 2026Updated 2 months ago
- [ICLR'26] Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology☆93Jan 26, 2026Updated 8 months ago
- ☆11Oct 4, 2018Updated 8 years ago
- An official codebase for "NormLens: Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Comm…☆10May 9, 2024Updated 2 years ago