AI & Computational Science

AI learns to search massive images by zooming and scanning like humans

How the science connects

Computer visionImage recognitionVisual attention

AI Insight

This paper introduces Visual Parallel Search (VPS), a framework that improves how AI models answer questions about high-resolution images by inspecting multiple image tiles simultaneously before selectively zooming in on relevant regions. The approach outperforms sequential zoom-only methods in 14 of 15 comparisons, with accuracy gains up to 8 points, particularly benefiting smaller models. The researchers also developed training methods using supervised fine-tuning and reinforcement learning that improved performance across multiple benchmarks while reducing unnecessary computational steps.


This advance could make visual question-answering systems more efficient and accurate for applications requiring analysis of detailed images, such as medical imaging, satellite imagery analysis, or document understanding. The method is especially promising for deployment scenarios where computational resources are limited, as it shows stronger improvements for smaller models.


Understand the Science

Computer vision 49 articles Explore Concept → Image recognition Concept coming soon Visual attention Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before obtaining a reliable overview. We introduce VPS, a visual parallel-search framework in which a main agent first invokes grid_search to inspect image tiles in parallel with question-conditioned sub-agents, and then adaptively invokes zoom_in PSisual Parallel Search improves mean accuracy over dedicated zoom-only search in 14 of 15 same-model comparisons, with gains up to 8.0 points and especially strong improvements for smaller main models. ZoomBench retains an approximately 3.2-point gain at every tested size. We further develop a supervision pipeline with hint-free verification and a paired role-specific GRPO surrogate for learning the controller and tile-reader roles. SFT improves observed accuracy on all five benchmark splits, including a 4.17-point gain on HR-Bench 4K. Role-specific RL further reshapes search behavior: main-only RL reduces mean tool use from 2.65 to 2.11 with similar pass@1 in an internal four-response evaluation, while external accuracy changes are mixed. Joint training reveals an asymmetry between local evidence reading and global search control. Together, these results support VPS as an effective inference-time scaffold and a trainable decomposition for visual evidence acquisition.

Source: Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom