HyperEyes is a parallel multimodal search agent that fuses visual grounding and retrieval into a single atomic action, enabling concurrent search across multiple entities while treating inference efficiency as a first-class training objective.
☆76May 23, 2026Updated 4 months ago
Alternatives and similar repositories for HyperEyes
Users that are interested in HyperEyes are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 🪐 Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback☆26Jan 29, 2026Updated 8 months ago
- Description for MV-MATH☆15Jul 20, 2025Updated last year
- The repository of VG-Refiner paper☆20Dec 9, 2025Updated 9 months ago
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning☆27Feb 8, 2026Updated 7 months ago
- Rewards as Labels: Revisiting RLVR from a Classification Perspective☆25Jun 26, 2026Updated 3 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆87May 8, 2026Updated 4 months ago
- The official code of "Towards Long-horizon Agentic Multimodal Search"☆30Apr 17, 2026Updated 5 months ago
- 🔍 OpenSearch-VL provides a fully open recipe for training strong multimodal deep search agents through high-quality data curation, diver…☆289Jul 30, 2026Updated 2 months ago
- [NeurIPS'25] The official code of "PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning"☆30Mar 30, 2026Updated 6 months ago
- [ICML 2026] ZwZ model family: SOTA fine-grained perception performace; ZoomBench: a new challenging perception benchmark☆191May 4, 2026Updated 4 months ago
- [CVPR2026] Official codebase for the paper "Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space"☆88May 12, 2026Updated 4 months ago
- Official data and code for the paper "VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents".☆14Mar 18, 2026Updated 6 months ago
- Official code, data, and models for "Hint Tuning: Less Data Makes Better Reasoners"☆22Jul 8, 2026Updated 2 months ago
- ☆48Jun 23, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆49Feb 9, 2026Updated 7 months ago
- ☆34Mar 17, 2026Updated 6 months ago
- [CVPR 2026] Official codes of "Monet: Reasoning in Latent Visual Space Beyond Image and Language"☆227Mar 19, 2026Updated 6 months ago
- [ICLR 2026] The official repository for the paper "AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning".☆84Aug 11, 2026Updated last month
- The implementation for ACL 2026: MatchTIR: Fine-Grained Supervision for Tool-Integrated Reasoning via Bipartite Matching.☆20Jul 27, 2026Updated 2 months ago
- Official codes of "Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs"☆20Sep 22, 2026Updated last week
- Benchmarking multimodal agents on realistic, ultra-challenging visual scenarios requiring long-horizon hybrid tool use.☆74Mar 10, 2026Updated 6 months ago
- ☆37Apr 13, 2026Updated 5 months ago
- Code and data for the paper: AI Sees Your Location—But With A Bias Toward The Wealthy World☆19Dec 15, 2025Updated 9 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- [CVPR 2026] Thinking with Programming Vision: Towards a Unified View for Thinking with Images☆74Jan 23, 2026Updated 8 months ago
- An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale☆635Updated this week
- Search Self-Play: Pushing the Frontier of Agent Capability without Supervision☆107Jul 23, 2026Updated 2 months ago
- Paper writing guide for Zhuang Liu Lab @ Princeton University☆37Jun 24, 2026Updated 3 months ago
- [NeurIPS 2025] Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing☆100Jul 27, 2025Updated last year
- [ICML 2026] XSkill: Continual Learning from Experience and Skills in Multimodal Agents☆271May 13, 2026Updated 4 months ago
- 学术主页 | Academic Page☆14Sep 24, 2026Updated last week
- ACL'2025: SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs. and preprint: SoftCoT++: Test-Time Scaling with Soft Chain-of…☆96May 30, 2025Updated last year
- https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT☆140Jan 30, 2026Updated 8 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆28Mar 17, 2026Updated 6 months ago
- Agentic MLLMs☆215Oct 24, 2025Updated 11 months ago
- ☆88Jun 16, 2026Updated 3 months ago
- [ECCV 26'] Official codebase for the paper LaViT☆35Jul 30, 2026Updated 2 months ago
- [CVPR 2026] OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe☆166Mar 30, 2026Updated 6 months ago
- Vision-OPD is a regional-to-global on-policy self-distillation framework that transfers a model's own privileged crop-conditioned percept…☆325Updated this week
- This repository contains the code for the paper “Neuro-Symbolic Query Compiler”, accepted to the Findings of ACL 2025.☆19Oct 20, 2025Updated 11 months ago