[ICLR 2026 Oral] Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
☆19Apr 29, 2026Updated 3 months ago
Alternatives and similar repositories for VC-STaR
Users that are interested in VC-STaR are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Proposed fuzzy reward model with GRPO to improve VLM's abilities in crowd counting task.☆21Apr 11, 2025Updated last year
- [NeurIPS 2025] This is the official repository for VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Se…☆15Oct 29, 2025Updated 9 months ago
- (CVPR25) Exploring Contextual Attribute Density in Referring Expression Counting☆20Dec 3, 2025Updated 8 months ago
- Implemented some industrial product surface defect detection using improved yolov5.☆12Apr 27, 2025Updated last year
- ☆28Feb 21, 2025Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- GRU-PPO for stable-baselines3.☆13Apr 24, 2024Updated 2 years ago
- PyTorch Implementation of the paper "MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training"☆26Jul 27, 2026Updated last week
- ☆11May 24, 2024Updated 2 years ago
- Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision☆47Oct 19, 2025Updated 9 months ago
- [CVPR 2024] S-DyRF: Reference-Based Stylized Radiance Fields for Dynamic Scenes☆13Jun 1, 2024Updated 2 years ago
- ☆17Dec 30, 2022Updated 3 years ago
- Domain generalization benchmark for skin lesion recognition, MICCAI 2023☆20Feb 13, 2024Updated 2 years ago
- MCP (Model Context Protocol) server for AFSIM simulation framework. Provides natural language scenario generation, entity management, sim…☆21May 20, 2026Updated 2 months ago
- [ACMMM 2026] PLUME: Latent Reasoning Based Universal Multimodal Embedding☆25Apr 29, 2026Updated 3 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ACCV 2022 Oral] SymmNeRF: Learning to Explore Symmetry Prior for Single-View View Synthesis☆14Mar 14, 2024Updated 2 years ago
- brain to speech☆13Mar 17, 2026Updated 4 months ago
- [MM 2023] Toward High Quality Facial Representation Learning☆20Oct 30, 2023Updated 2 years ago
- ☆12Feb 24, 2023Updated 3 years ago
- Official repository for the paper "Towards Interpretable Counterfactual Generation via Multimodal Autoregression"☆17Nov 7, 2025Updated 8 months ago
- (CVPR 26) Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration☆37Mar 8, 2026Updated 4 months ago
- Official repository for "SODA: Bottleneck Diffusion Models for Representation Learning"☆28Mar 21, 2024Updated 2 years ago
- Pipelined MIPS architecture created in Verilog. Includes data forwarding and hazard detection.☆16Apr 1, 2018Updated 8 years ago
- Deep Counterfactual Prediction with Categorical Backward Variables☆12Feb 8, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICLR 2026 Oral & ICML 2026] Generative Universal Verifier as Multimodal Meta-Reasoner☆64May 29, 2026Updated 2 months ago
- [ACL 2025] "CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought"☆17Apr 3, 2025Updated last year
- Mental image reconstruction from human brain activity☆17Jul 1, 2024Updated 2 years ago
- ☆25Mar 15, 2023Updated 3 years ago
- Official Repository for NeurIPS'25 Paper "Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task"☆23May 18, 2026Updated 2 months ago
- This is a PyTorch implementation of 3DRefTR proposed by our paper "A Unified Framework for 3D Point Cloud Visual Grounding"☆26Aug 24, 2023Updated 2 years ago
- 🔥This is a repository of paper list for streaming LLMs/MLLMs.☆26Apr 19, 2026Updated 3 months ago
- Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT☆14Jul 30, 2025Updated last year
- [CVPR 2025] DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval☆22Jun 23, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- The PyTorch implementation for "DEAL: Disentangle and Localize Concept-level Explanations for VLMs" (ECCV 2024 Strong Double Blind)☆20Mar 9, 2026Updated 4 months ago
- Code for "General-Purpose Brain Foundation Models for Time-Series Neuroimaging Data"☆15Dec 14, 2024Updated last year
- Python script to obtain dynamic functional connectivity metrics, after using a sliding window approach, statistical analyses to test for …☆12Sep 10, 2024Updated last year
- chinesetokenization☆13Jun 4, 2013Updated 13 years ago
- 一个小小的书单,收集整理了一些计算机科学与技术方面的书籍英文原著pdf。☆10Jan 13, 2022Updated 4 years ago
- [ICLR 2026] Official code for paper: TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinf…☆27Jan 29, 2026Updated 6 months ago
- We introduce XBrainLab, an open-source user-friendly software, for accelerated interpretation of neural patterns from EEG data based on c…☆14Dec 5, 2025Updated 7 months ago