☆17Apr 15, 2026Updated 5 months ago
Alternatives and similar repositories for Q-Zoom
Users that are interested in Q-Zoom are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR2026] Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception☆18Jan 26, 2026Updated 7 months ago
- [NeurIPS 2025] MODEM: A Morton-Order Degradation Estimation Mechanism for Adverse Weather Image Recovery.☆24May 17, 2026Updated 3 months ago
- Official project page and code repository for WiT, a pixel space diffusion☆17Updated this week
- Official pytorch implementation of EMNLP 2022 long paper “A Sequential Flow Control Framework for Multi-hop Knowledge Base Question Answe…☆15Apr 22, 2023Updated 3 years ago
- [AAAI 2025] Efficient Image-to-Image Diffusion Classifier for Adversarial Robustness☆20Aug 21, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICLR 2026] VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models☆21Feb 22, 2026Updated 6 months ago
- Action aware Dynamic Pruning for Efficient Vision Language Action Manipulation☆27Feb 28, 2026Updated 6 months ago
- DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference☆25May 21, 2026Updated 3 months ago
- The offical code of PolarBEV (CoRL2022).☆58Sep 17, 2022Updated 3 years ago
- minimal library for image super-resolution implemented in jax☆13Jul 12, 2021Updated 5 years ago
- documentation website for SimpleFOCproject☆17Apr 9, 2026Updated 5 months ago
- PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding☆31Jun 10, 2026Updated 3 months ago
- Jointly Optimizing Large Language Models for Reasoning and Self-Refinement☆14Updated this week
- Offical implementation for "Trash or Treasure? An Interactive Dual-Stream Strategy for Single Image Reflection Separation".☆63Aug 22, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [CVPR2026] Official codebase for the paper "Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space"☆88May 12, 2026Updated 4 months ago
- Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement☆21Jul 4, 2026Updated 2 months ago
- Squeeze verbose LLM agent tool output down to only the relevant lines☆23Apr 27, 2026Updated 4 months ago
- ☆15Jul 18, 2026Updated last month
- Official implementation for "Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation Decoders"☆27Aug 19, 2026Updated 3 weeks ago
- From Local Matches to Global Masks: Novel Instance Detection in Open-World Scenes☆26Jul 28, 2026Updated last month
- IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning☆18Feb 27, 2026Updated 6 months ago
- ☆14Sep 22, 2025Updated 11 months ago
- [CVPRW 2026 Oral] Less Detail, Better Answers: Degradation-Driven Prompting for VQA☆20Apr 25, 2026Updated 4 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official Pytorch Code for "Rethinking Degradation: Radiograph Super-Resolution via AID-SRGAN" - MICCAI 2022 Workshop☆16Dec 11, 2024Updated last year
- [WACV 2026] ZonUI-3B — A lightweight, resolution-aware GUI grounding model trained with only 24K samples on a single RTX 4090.☆26Jan 2, 2026Updated 8 months ago
- ☆22Jul 30, 2026Updated last month
- MCOUT: Multimodal Chain of Continuous Thought for Latent Reasoning☆22Oct 4, 2025Updated 11 months ago
- AuthFace: Towards Authentic Blind Face Restoration with Face-oriented Generative Diffusion Prior (ACM MM 2025 Oral)☆20Mar 5, 2026Updated 6 months ago
- CVPR25☆28Jul 2, 2025Updated last year
- ☆14Jan 13, 2025Updated last year
- ICRA2026: ABPolicy Asynchronous B-Spline Flow Policy for Real-Time and Smooth Robotic Manipulation☆48Apr 22, 2026Updated 4 months ago
- A simple visual test-time scaling method for GUI agent grounding☆26Dec 7, 2025Updated 9 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- GPT Table Semantic Parsing with complex & non-intuitive structure.☆17Jul 16, 2025Updated last year
- A Practical Zoom-in GUI Grounding and Behavior-Based Evaluation method.☆27Aug 27, 2026Updated 2 weeks ago
- Official codes of "Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs"☆19Feb 15, 2026Updated 7 months ago
- ☆20Jul 29, 2025Updated last year
- [EMNLP Main 2026]VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning.☆26Jul 20, 2026Updated last month
- An arbitrage bot is a smart contract connected to an external automation script that controls its operation.☆2,699Updated this week
- OTSeq2Set, XMTC☆11Dec 31, 2022Updated 3 years ago