[ICLR 2026 Oral] Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
☆21Apr 29, 2026Updated 3 months ago
Alternatives and similar repositories for VC-STaR
Users that are interested in VC-STaR are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- (CVPR25) Exploring Contextual Attribute Density in Referring Expression Counting☆20Dec 3, 2025Updated 8 months ago
- A large-scale training and benchmarking framework for rPPG.☆10Nov 26, 2024Updated last year
- ☆12Jan 6, 2025Updated last year
- VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models☆79Jul 13, 2024Updated 2 years ago
- ☆14Nov 26, 2025Updated 8 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- (AAAI 2024) Paper: Semi-Supervised Class-Agnostic Motion Prediction with Pseudo Label Regeneration and BEVMix☆14Sep 10, 2024Updated last year
- ☆28Feb 21, 2025Updated last year
- PyTorch Implementation of the paper "MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training"☆26Aug 16, 2026Updated last week
- [EMNLP'2023 Findings] MoqaGPT, for zero-shot multimodal question answering with LLMs☆13Dec 28, 2024Updated last year
- ☆11May 24, 2024Updated 2 years ago
- Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision☆47Oct 19, 2025Updated 10 months ago
- [CVPR 2024] S-DyRF: Reference-Based Stylized Radiance Fields for Dynamic Scenes☆13Jun 1, 2024Updated 2 years ago
- The official implementation of the paper "Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study"☆17May 8, 2026Updated 3 months ago
- ☆16Apr 30, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Domain generalization benchmark for skin lesion recognition, MICCAI 2023☆20Feb 13, 2024Updated 2 years ago
- A curated list of all awesome pygames created by Agneay B Nair☆12Apr 28, 2024Updated 2 years ago
- [ACMMM 2026] PLUME: Latent Reasoning Based Universal Multimodal Embedding☆25Apr 29, 2026Updated 3 months ago
- [ACCV 2022 Oral] SymmNeRF: Learning to Explore Symmetry Prior for Single-View View Synthesis☆14Mar 14, 2024Updated 2 years ago
- ☆12Feb 24, 2023Updated 3 years ago
- Official repository for the paper "Towards Interpretable Counterfactual Generation via Multimodal Autoregression"☆17Nov 7, 2025Updated 9 months ago
- Official repository for "SODA: Bottleneck Diffusion Models for Representation Learning"☆28Mar 21, 2024Updated 2 years ago
- Pipelined MIPS architecture created in Verilog. Includes data forwarding and hazard detection.☆16Apr 1, 2018Updated 8 years ago
- Deep Counterfactual Prediction with Categorical Backward Variables☆12Feb 8, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A new baseline and benchmark for robust monocular depth estimation☆23Feb 11, 2025Updated last year
- Mental image reconstruction from human brain activity☆17Jul 1, 2024Updated 2 years ago
- ☆25Mar 15, 2023Updated 3 years ago
- Official Repository for NeurIPS'25 Paper "Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task"☆23May 18, 2026Updated 3 months ago
- 🔥This is a repository of paper list for streaming LLMs/MLLMs.☆27Apr 19, 2026Updated 4 months ago
- Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT☆14Jul 30, 2025Updated last year
- [ICLR 2026] The official repository for paper "ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning"☆192May 1, 2026Updated 3 months ago
- The PyTorch implementation for "DEAL: Disentangle and Localize Concept-level Explanations for VLMs" (ECCV 2024 Strong Double Blind)☆20Mar 9, 2026Updated 5 months ago
- Code for "General-Purpose Brain Foundation Models for Time-Series Neuroimaging Data"☆15Dec 14, 2024Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Python script to obtain dynamic functional connectivity metrics, after using a sliding window approach, statistical analyses to test for …☆12Sep 10, 2024Updated last year
- Official Implement of the paper "Unifying Segment Anything in Microscopy with Multimodal Large Language Model"☆20Apr 27, 2026Updated 3 months ago
- 一个小小的书单,收集整理了一些计算机科学与技术方面的书籍英文原著pdf。☆10Jan 13, 2022Updated 4 years ago
- Traffic Video Event Retrieval via Text Query using Vehicle Appearance and Motion Attributes☆10May 31, 2026Updated 2 months ago
- [ICLR 2026] Official code for paper: TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinf…☆27Jan 29, 2026Updated 6 months ago
- ATM-Bench: A benchmark for long-term personalized memory QA spanning ~4 years of multimodal data (images, videos, emails). Features refer…☆61Aug 13, 2026Updated last week
- LeaderBoard for various CBIR models☆24Feb 5, 2018Updated 8 years ago