v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
☆21Aug 10, 2026Updated 3 weeks ago
Alternatives and similar repositories for v1
Users that are interested in v1 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Jax implementation of VIT-VQGAN☆10Jan 25, 2024Updated 2 years ago
- [NeurIPS 2025] This is the official repository for VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Se…☆15Oct 29, 2025Updated 10 months ago
- ☆19Nov 8, 2025Updated 9 months ago
- Corpus to accompany: "Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding"☆11Apr 11, 2025Updated last year
- ☆30Mar 13, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This is the implementation of CounterCurate, the data curation pipeline of both physical and semantic counterfactual image-caption pairs.☆19Jun 27, 2024Updated 2 years ago
- [ICCV 2025] Auto Interpretation Pipeline and many other functionalities for Multimodal SAE Analysis.☆200Sep 26, 2025Updated 11 months ago
- Download Web-10K data by querying Bing Image Search☆10Feb 1, 2022Updated 4 years ago
- Ultra-minimal autoregressive diffusion model for image generation☆21Jul 20, 2026Updated last month
- Course materials for Aircraft Dynamics (ASEN 3728) at CU Boulder☆14May 2, 2025Updated last year
- [CVPR 2026] Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens☆297Aug 2, 2025Updated last year
- VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning☆37May 9, 2026Updated 3 months ago
- ☆12Dec 9, 2025Updated 8 months ago
- ☆15Jan 9, 2026Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Train vector quantized CLIP models using pytorch lightning☆21Jul 14, 2024Updated 2 years ago
- ☆15Apr 8, 2022Updated 4 years ago
- ☆12Jun 18, 2024Updated 2 years ago
- [AAAI'25]: Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP☆23Aug 5, 2025Updated last year
- ShopEasy is a containerized e‑commerce boilerplate with an Angular SPA and Spring Boot API (JWT‑secure), backed by MySQL and Kafka, deplo…☆17Sep 5, 2025Updated 11 months ago
- [NeurIPS 2025] MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning☆107Sep 19, 2025Updated 11 months ago
- ☆78Apr 9, 2026Updated 4 months ago
- Code for EMNLP2021 paper “Transductive Learning for Unsupervised Text Style Transfer”☆12Sep 19, 2021Updated 4 years ago
- A general purpose gpu computation library in odin☆12May 9, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [NeurIPS 2024] Calibrated Self-Rewarding Vision Language Models☆87Oct 26, 2025Updated 10 months ago
- This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception o…☆29Jul 9, 2025Updated last year
- Code for paper 'Zero-Shot Scene Graph Generation via Triplet Calibration and Reduction' (TOMM 2023)☆10Sep 6, 2025Updated 11 months ago
- In this codebase we establish a benchmark for egocentric user adaptation based on Ego4d.First, we start from a population model which ha…☆15Jul 24, 2026Updated last month
- ☆14Jun 12, 2024Updated 2 years ago
- [not maintained anymore] [for study purpose] A simple PyTorch implementation for "Global Vectors for Word Representation".☆17Nov 7, 2019Updated 6 years ago
- [ICLR'26] Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology☆93Jan 26, 2026Updated 7 months ago
- An official codebase for "NormLens: Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Comm…☆10May 9, 2024Updated 2 years ago
- ☆12Dec 20, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Code accompanying the paper "Noise Contrastive Alignment of Language Models with Explicit Rewards" (NeurIPS 2024)☆59Nov 8, 2024Updated last year
- An introduction course about scientific computing and research for the first year undergraduates taught at Fudan University.☆10Nov 6, 2025Updated 9 months ago
- A collection of tools that make creating and consuming callbags more intuitive☆12Dec 1, 2025Updated 8 months ago
- Building and Deploying a Secure ReactJS App with Docker, NGINX, and Automating with GitHub Actions to AWS EC2☆10Mar 26, 2025Updated last year
- Reasoning in Large Language Models: Papers and Resources, including Chain-of-Thought and OpenAI o1 🍓☆20Oct 10, 2024Updated last year
- Compile Markdown files to beautiful PDF documents by pandoc and tectonic.☆11Jul 6, 2026Updated last month
- ☆13Feb 16, 2024Updated 2 years ago