[ICLR 2026 Oral] Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
☆23Apr 29, 2026Updated 4 months ago
Alternatives and similar repositories for VC-STaR
Users that are interested in VC-STaR are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2025] This is the official repository for VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Se…☆15Oct 29, 2025Updated 10 months ago
- Proposed fuzzy reward model with GRPO to improve VLM's abilities in crowd counting task.☆21Apr 11, 2025Updated last year
- (CVPR25) Exploring Contextual Attribute Density in Referring Expression Counting☆20Dec 3, 2025Updated 9 months ago
- A large-scale training and benchmarking framework for rPPG.☆10Nov 26, 2024Updated last year
- VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models☆78Jul 13, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Implemented some industrial product surface defect detection using improved yolov5.☆12Apr 27, 2025Updated last year
- ☆28Feb 21, 2025Updated last year
- This is the official code for the paper "Reconstruct before Query: Continual Missing Modality Learning with Decomposed Prompt Collaborati…☆12Aug 13, 2024Updated 2 years ago
- PyTorch Implementation of the paper "MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training"☆26Aug 29, 2026Updated 2 weeks ago
- [EMNLP'2023 Findings] MoqaGPT, for zero-shot multimodal question answering with LLMs☆13Dec 28, 2024Updated last year
- ☆11May 24, 2024Updated 2 years ago
- Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision☆47Oct 19, 2025Updated 10 months ago
- The official implementation of the paper "Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study"☆18May 8, 2026Updated 4 months ago
- ☆16Apr 30, 2026Updated 4 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆17Dec 30, 2022Updated 3 years ago
- Domain generalization benchmark for skin lesion recognition, MICCAI 2023☆20Feb 13, 2024Updated 2 years ago
- A curated list of all awesome pygames created by Agneay B Nair☆12Apr 28, 2024Updated 2 years ago
- [ACMMM 2026] PLUME: Latent Reasoning Based Universal Multimodal Embedding☆25Apr 29, 2026Updated 4 months ago
- [MM 2023] Toward High Quality Facial Representation Learning☆19Oct 30, 2023Updated 2 years ago
- PyTorch implementation of Graph Convolutional Networks in Feature Space for Image Deblurring and Super-resolution, IJCNN 2021.☆12Nov 14, 2021Updated 4 years ago
- Redundancy Undermines the Trustworthiness of Self-Interpretable GNNs, International Conference on Machine Learning (ICML), 2025☆15Jun 23, 2025Updated last year
- Official repository for the paper "Towards Interpretable Counterfactual Generation via Multimodal Autoregression"☆17Nov 7, 2025Updated 10 months ago
- (CVPR 26) Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration☆41Mar 8, 2026Updated 6 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Official repository for "SODA: Bottleneck Diffusion Models for Representation Learning"☆28Mar 21, 2024Updated 2 years ago
- Pipelined MIPS architecture created in Verilog. Includes data forwarding and hazard detection.☆16Apr 1, 2018Updated 8 years ago
- Deep Counterfactual Prediction with Categorical Backward Variables☆12Feb 8, 2023Updated 3 years ago
- [ICLR 2026 Oral & ICML 2026] Generative Universal Verifier as Multimodal Meta-Reasoner☆67May 29, 2026Updated 3 months ago
- [ACL 2025] "CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought"☆17Apr 3, 2025Updated last year
- Mental image reconstruction from human brain activity☆18Jul 1, 2024Updated 2 years ago
- ☆25Mar 15, 2023Updated 3 years ago
- Official Repository for NeurIPS'25 Paper "Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task"☆23Sep 4, 2026Updated last week
- Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT☆14Jul 30, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [ICLR 2026] The official repository for paper "ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning"☆198May 1, 2026Updated 4 months ago
- [CVPR 2025] DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval☆22Jun 23, 2025Updated last year
- The PyTorch implementation for "DEAL: Disentangle and Localize Concept-level Explanations for VLMs" (ECCV 2024 Strong Double Blind)☆20Mar 9, 2026Updated 6 months ago
- Code for "General-Purpose Brain Foundation Models for Time-Series Neuroimaging Data"☆15Dec 14, 2024Updated last year
- chinesetokenization☆13Jun 4, 2013Updated 13 years ago
- 一个小小的书单,收集整理了一些计算机科学与技术方面的书籍英文原著pdf。☆10Jan 13, 2022Updated 4 years ago
- We introduce XBrainLab, an open-source user-friendly software, for accelerated interpretation of neural patterns from EEG data based on c…☆14Dec 5, 2025Updated 9 months ago