Remote Sensing Novel View Synthesis With Implicit Multiplane Representations

IEEE Transactions on Geoscience and Remote Sensing2022

Yongchang Wu Zhengxia Zou* Zhenwei Shi

School of Astronautics, Beihang University

* Corresponding author

Sparse input views initialize and optimize an implicit multiplane representation; green cameras mark training views and red cameras mark novel views rendered by ImMPI

ImMPI combines learned scene priors and implicit multiplane representations to reconstruct remote sensing scenes from sparse posed images and synthesize novel views through differentiable rendering.

Abstract

Novel view synthesis of remote sensing (RS) scenes is of great significance for scene visualization, human–computer interaction, and various downstream applications. Despite the recent advances in computer graphics and photogrammetry technology, generating novel views is still challenging particularly for RS images due to its high complexity, view sparsity, and limited view-perspective variations. In this article, we propose a novel RS view synthesis method by leveraging the recent advances in implicit neural representations. Considering the overhead and far depth imaging of RS images, we represent the 3-D space by combining implicit multiplane images (ImMPI) representation and deep neural networks. The 3-D scene is reconstructed under a self-supervised optimization paradigm through a differentiable multiplane renderer with multiview input constraints. Images from any novel views thus can be freely rendered on the basis of the reconstructed model. As a by-product, the depth maps corresponding to the given viewpoint can be generated along with the rendering output. We refer to our method as ImMPIs. To further improve the view synthesis under sparse-view inputs, we explore the learning-based initialization of RS 3-D scenes and proposed a neural-network-based prior extractor to accelerate the optimization process. In addition, we propose a new dataset for RS novel view synthesis with multiview real-world Google Earth images. Extensive experiments demonstrate the superiority of the ImMPI over previous state-of-the-art methods in terms of reconstruction accuracy, visual fidelity, and time efficiency. Ablation experiments also suggest the effectiveness of our methodology design.

Method

ImMPI learns scene priors, refines each scene’s implicit representation, and renders novel views from explicit planes.

Two-stage ImMPI pipeline: cross-scene initialization with a prior extractor, followed by per-scene optimization through differentiable rendering
The prior extractor initializes the MPI generator from a reference view. Multiview supervision then refines the generator while the prior extractor stays fixed.
01

Learn across scenes

Learn scene priors from posed RGB images.

02

Optimize per scene

Refine the MPI generator using the scene’s input views.

03

Render new viewpoints

Warp and composite planes to render novel views and depth.

Results

Novel-view synthesis

ImMPI achieves higher average image quality and sharper novel-view renderings than NeRF and NeRF++ on held-out LEVIR-NVS viewpoints.

Train and test PSNR, SSIM and LPIPS for NeRF, NeRF++ and ImMPI across all 16 LEVIR-NVS scenes
Quantitative comparison on LEVIR-NVS. Each pair reports train-view / test-view performance.
First scene, two unseen viewpoints. Columns: scene, NeRF, NeRF++, ImMPI, and ground truth.
Second scene, two unseen viewpoints. Columns: scene, NeRF, NeRF++, ImMPI, and ground truth.
Third scene, two unseen viewpoints. Columns: scene, NeRF, NeRF++, ImMPI, and ground truth.
Fourth scene, two unseen viewpoints. Columns: scene, NeRF, NeRF++, ImMPI, and ground truth.
Qualitative comparison at test viewpoints excluded from pretraining and per-scene optimization. Columns: scene, NeRF, NeRF++, ImMPI, and ground truth.

Inside the multiplane representation

Different depth planes capture different parts of the scene, from rooftops to the ground. Explore the color and sigma values produced by the optimized representation.

These visualizations illustrate the representation; they are not depth-accuracy measurements.

Observation · Sigma values separate scene content across depth planes.

Observation · RGB values of the explicit MPI layers, alongside the rendered image and depth.

Depth as a by-product

The same planes can render a depth map without training on ground-truth depth. Multiview refinement produces more detailed structure than initialization alone.

Errors remain where depth changes sharply, including object boundaries. The paper’s depth evidence is qualitative; the image-quality results do not establish metric depth or surveying accuracy.

Depth comparison for the first pair of scenes. Columns: RGB image, depth without per-scene optimization, optimized depth, and ground truth.
Depth comparison for the second pair of scenes. Columns: RGB image, depth without per-scene optimization, optimized depth, and ground truth.
Cool colors indicate points closer to the camera; warm colors indicate points farther away.

Dataset

LEVIR-NVS contains sixteen scenes spanning buildings, urban areas, villages and mountains. The dataset uses 3D scene models captured from Google Earth, with multiview images generated in Blender.

Images per scene
21
Image resolution
512 × 512
Train / test views
11 / 10
Dataset and camera files
The 16-scene collection. Posed RGB images are used for per-scene optimization and evaluation.

Citation

@ARTICLE{9852475,
  author={Wu, Yongchang and Zou, Zhengxia and Shi, Zhenwei},
  journal={IEEE Transactions on Geoscience and Remote Sensing},
  title={Remote Sensing Novel View Synthesis With Implicit Multiplane Representations},
  year={2022},
  volume={60},
  number={},
  pages={1-13},
  doi={10.1109/TGRS.2022.3197409}}

Published in IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–13, 2022. DOI: 10.1109/TGRS.2022.3197409

Figure

Open image