Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
☆19Jul 4, 2026Updated 2 weeks ago
Alternatives and similar repositories for Zoom-Refine
Users that are interested in Zoom-Refine are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PyTorch Implementation of "Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Larg…☆49Mar 2, 2026Updated 4 months ago
- Official PyTorch implementation of `[ACMMM 2023]Relational Contrastive Learning for Scene Text Recognition`☆17Sep 22, 2023Updated 2 years ago
- 💻 哈工大硕士生《数值分析》课程上机 实验代码,看看有没有你需要的☆17Dec 20, 2021Updated 4 years ago
- ☆15Apr 15, 2026Updated 3 months ago
- [ICLR 2026] Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models☆29Mar 21, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆27Jul 20, 2024Updated 2 years ago
- Unofficial implementation of the Ask-LLM paper 'How to Train Data-Efficient LLMs', arXiv:2402.09668.☆12Jun 19, 2024Updated 2 years ago
- Code for Retrieval-Augmented Perception (ICML 2025)☆73Apr 22, 2026Updated 2 months ago
- ☆10Jul 11, 2022Updated 4 years ago
- [NeurIPS 2024] Official code for the paper 'RankUp: Boosting Semi-Supervised Regression with an Auxiliary Ranking Classifier'☆14Aug 22, 2025Updated 10 months ago
- ✨✨ [ICLR 2025] MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?☆160Oct 21, 2025Updated 8 months ago
- Official repo for "TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders"☆24Apr 9, 2026Updated 3 months ago
- ☆14Apr 9, 2026Updated 3 months ago
- ☆13Jan 14, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [WACV 2026] ZonUI-3B — A lightweight, resolution-aware GUI grounding model trained with only 24K samples on a single RTX 4090.☆26Jan 2, 2026Updated 6 months ago
- [ICLR'26] Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology☆92Jan 26, 2026Updated 5 months ago
- [ACL2026 Findings] "Towards Hierarchical Multi-Step Reward Models for Enhanced Reasoning in Large Language Models"☆20Mar 25, 2025Updated last year
- (ICML 2025) Rethinking Chain-of-Thought from the Perspective of Self-Training☆13Feb 15, 2025Updated last year
- [ACM MM 2025] ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models☆18Jul 15, 2025Updated last year
- ICCV'23 | Adverse Weather Removal with Codebook Priors☆10Aug 28, 2023Updated 2 years ago
- A simple visual test-time scaling method for GUI agent grounding☆26Dec 7, 2025Updated 7 months ago
- ☆13Jun 21, 2025Updated last year
- [EMNLP 2024] Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality☆22Oct 8, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official implementation of RIS-FUSION (ICASSP 2026 Oral)☆21Mar 13, 2026Updated 4 months ago
- ☆27Feb 3, 2026Updated 5 months ago
- (CVPR2024) MeaCap: Memory-Augmented Zero-shot Image Captioning☆56Aug 16, 2024Updated last year
- ☆18Feb 18, 2025Updated last year
- ☆12Feb 19, 2024Updated 2 years ago
- 首届阿里云弹性计算挑战赛云资源调度赛道作品☆15Mar 31, 2021Updated 5 years ago
- Official Repository for NeurIPS'25 Paper "Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task"☆23May 18, 2026Updated 2 months ago
- [EMNLP 2025]Repository for paper "DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning"☆30Jul 2, 2025Updated last year
- ☆18Dec 23, 2025Updated 6 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- [CCS-LAMPS'24] LLM IP Protection Against Model Merging☆16Oct 14, 2024Updated last year
- CVPR 2022 (Oral) Pytorch Code for Unsupervised Vision-and-Language Pre-training via Retrieval-based Multi-Granular Alignment☆21Apr 15, 2022Updated 4 years ago
- ☆14Oct 31, 2022Updated 3 years ago
- ☆20Jun 17, 2024Updated 2 years ago
- [ICLR2025] HiLo: A Learning Framework for Generalized Category Discovery Robust to Domain Shifts☆22Aug 1, 2025Updated 11 months ago
- ☆59Feb 20, 2022Updated 4 years ago
- ☆14Sep 14, 2021Updated 4 years ago