Physics

Physically Grounded Monocular Depth via Nanophotonic Wavefront Encoding

How the science connects

Computer visionNanophotonicsDepth perception

AI Insight

Researchers have developed a method to improve monocular depth estimation by integrating metalenses—ultrathin optical elements—with depth foundation models. The metalenses physically encode depth information by creating depth-dependent positional shifts in two polarized light wavefronts, which are then processed by a fine-tuned pretrained depth model. The team created a simulation pipeline to generate training data from RGB-D datasets that accounts for physical factors, and their experiments show the approach outperforms existing monocular depth estimation and depth-from-defocus methods.


This technology could enable more accurate 3D perception for applications in robotics, autonomous vehicles, and augmented reality using single-camera systems. By physically encoding depth cues at the optical level rather than relying solely on computational inference, the method addresses a fundamental limitation in monocular depth estimation regarding metric scale ambiguity.


Understand the Science

Computer vision 45 articles Explore Concept → Nanophotonics Concept coming soon Depth perception Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Depth foundation models (DFMs) offer strong learned priors for 3D perception from single RGB images but lack physical depth cues, leading to ambiguities in metric scale. We introduce metalenses, an emerging class of ultrathin planar optical elements, as a solution to physically encode missing metric depth cues via nanophotonics. In this paper, we bridge the gap between metalens and DFMs to achieve accurate metric monocular depth sensing. In a single monocular shot, our metalens embeds depth-dependent positional shifts into two polarized optical wavefronts. With an input adaptation strategty, we enable direct fine-tuning that aligns a pretrained DFM with the optical signals. To scale the training data, we further develop a comprehensive simulation pipeline that synthesizes metalens responses from RGB-D datasets, incorporating physical factors to minimize the sim-to-real gap. Experiments demonstrate that this approach outperforms both monocular metric depth estimation and depth-from-defocus baselines, showing an effective pathway for accurate monocular metric depth sensing.

Source: Physically Grounded Monocular Depth via Nanophotonic Wavefront Encoding