AI Insight
AnyView is a new diffusion-based video generation framework that synthesizes novel viewpoints in dynamic scenes without requiring strict geometric assumptions. The system is trained on diverse datasets (2D monocular, 3D multi-view static, and 4D multi-view dynamic) to create a generalist model capable of generating spatiotemporally consistent videos from arbitrary camera positions and trajectories. Unlike existing methods that require significant viewpoint overlap, AnyView maintains video quality and consistency even when generating views from extreme camera angles in real-world scenarios.
Why it matters
This technology could significantly advance applications in virtual reality, autonomous vehicle simulation, film production, and robotics by enabling realistic video synthesis from any desired viewpoint in dynamic environments. The introduction of AnyViewBench as a challenging benchmark also provides the research community with a new standard for evaluating extreme dynamic view synthesis capabilities.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
-cross
Abstract: Modern generative video models excel at producing convincing, high-quality outputs, but struggle to maintain multi-view and spatiotemporal consistency in highly dynamic real-world environments. In this work, we introduce $textbf{AnyView}$, a diffusion-based video generation framework for $textit{dynamic view synthesis}$ with minimal inductive biases or geometric assumptions. We leverage multiple data sources with various levels of supervision, including monocular (2D), multi-view static (3D) and multi-view dynamic (4D) datasets, to train a generalist spatiotemporal implicit representation capable of producing zero-shot novel videos from arbitrary camera locations and trajectories. We evaluate AnyView on standard benchmarks, showing competitive results with the current state of the art, and propose $textbf{AnyViewBench}$, a challenging new benchmark tailored towards $textit{extreme}$ dynamic view synthesis in diverse real-world scenarios. In this more dramatic setting, we find that most baselines drastically degrade in performance, as they require significant overlap between viewpoints, while AnyView maintains the ability to produce realistic, plausible, and spatiotemporally consistent videos when prompted from $textit{any}$ viewpoint. Results, data, code, and models can be viewed at: https://tri-ml.github.io/AnyView/
Source: AnyView: Synthesizing Any Novel View in Dynamic Scenes