Progressive Gaussian Splatting for High-Fidelity 3D Reconstruction of Large Scale Remote Sensing Scenes

IEEE Geoscience and Remote Sensing Magazine2026

Yongchang WuZhenwei ShiZhengxia Zou✉

Department of Aerospace Intelligent Science and Technology, School of Astronautics, Beihang University

State Key Laboratory of Virtual Reality Technology and Systems, Beihang University

Corresponding author: Zhengxia Zou

Sequential UAV images

A sequence of aerial observations acquired along the UAV flight path

Incremental Gaussian reconstruction

A persistent Gaussian map with an active local region updated from incoming observations
ProAGS incrementally reconstructs large-scale 3D scenes from posed UAV image sequences, expanding the map while preserving earlier reconstructions.

Abstract

Unmanned Aerial Vehicles (UAVs) have become essential platforms for large-scale remote sensing, enabling rapid and high-resolution data acquisition. Achieving efficient, high-fidelity 3D large-scale scene reconstruction and novel view synthesis from streaming UAV imagery is of critical importance for many applications. Recent advances in 3D Gaussian Splatting (3DGS) provide an efficient and high-fidelity representation for novel view synthesis, but existing 3DGS pipelines are not tailored to the sequential, nadir-looking characteristics of UAV imagery and struggle to scale to kilometer-scale scenes. In this paper, we present a progressive aerial-oriented Gaussian splatting reconstruction framework (ProAGS). To address the depth ambiguity caused by large imaging distances and limited inter-frame disparity, we introduce a multi-view dense depth estimation module that provides scale-consistent and reliable depth priors for Gaussian initialization and geometric supervision. We further reinterpret progressive mapping as a continual learning problem and propose a loss-aware replay strategy together with a per-Gaussian learning-rate scheduler, which jointly mitigate catastrophic forgetting in previously mapped regions and overfitting in newly initialized primitives. To achieve scalability, we design a bird’s-eye-view (BEV) grid structure and a dynamic CPU–GPU map loading mechanism that decouple GPU memory usage from scene extent and support progressive mapping over large-scale UAV scenes. Experiments on challenging aerial datasets demonstrate that the proposed method produces more stable geometry, fewer artifacts, and competitive rendering quality compared with existing incremental 3DGS baselines, while maintaining practical runtime and memory footprints for UAV remote sensing applications.

Method

ProAGS builds a persistent Gaussian map from sequential UAV images through two asynchronous threads for image processing and progressive mapping. Depth-guided initialization adds reliable geometry from multi-view depth; loss-aware replay and per-Gaussian learning rates preserve earlier views while adapting to new observations; and BEV map management keeps the global map in CPU memory while loading active regions onto the GPU.

ProAGS framework: keyframe and neighbor selection, MVS depth estimation, progressive mapping, and CPU–GPU BEV map management
Sequential image processing supplies depth and uncertainty; progressive mapping updates the scene; BEV cells organize the global and active local maps.

Select and estimate

Select keyframes that reveal new regions. Retrieve co-visible neighbors with a useful baseline and estimate dense depth using DiffMVS.

Expand and optimize

Initialize reliable new geometry, replay past views according to their reconstruction loss, and schedule each Gaussian’s learning rate independently.

Load and retain

Keep the global BEV map in CPU memory. Transfer the local cells needed for current and recent views to the GPU.

Depth-guided initialization

New regions are identified by low rendered opaqueness or disagreement between rendered and estimated depth. Confidence filtering rejects unreliable MVS depth before back-projection. Observation frames refine the existing map without initializing new geometry.

Depth and opaqueness masks identify newly observed regions for reliable Gaussian initialization
The union of the two masks identifies candidate regions for new primitives.
Loss-aware replay and per-Gaussian learning rates

Loss-aware replay revisits past views according to their reconstruction loss, helping preserve previously mapped regions.

Each Gaussian follows its own learning-rate warmup and decay, balancing adaptation of new geometry with stability of the existing map.

Loss-weighted view replay and a per-Gaussian warmup followed by learning-rate decay
A separate optimization history for each primitive.

An optional global refinement stage optimizes the completed map for 10,000 additional iterations. Results with this stage are denoted ProAGS*.

Results

We evaluate image quality across seven UAV scenes with PSNR, SSIM, and LPIPS, alongside reconstructed-view comparisons. We also examine processing time, memory use, and retention of previously observed views.

Quantitative comparison

ProAGS improves LPIPS over VINGS-Mono on all seven scenes. PSNR and SSIM gains are not uniform: Countryside and Mountain have lower PSNR, and Mountain has lower SSIM. Optional global refinement improves all three metrics over ProAGS in every scene.

PSNR, SSIM, and LPIPS for five methods across Factory, Countryside, Farmland, Parking, Town, Mountain, and Suburb
PSNR (dB) and SSIM ↑; LPIPS ↓. Red/blue denote best/second-best results among MonoGS*, VINGS-Mono, and ProAGS. BlockGaussian uses global optimization. All methods receive offline poses; MonoGS* also uses DiffMVS initialization and scene partitioning. ProAGS* adds global refinement.

Quantitative comparison

Ground-truth photographs and reconstructed views across Factory, Countryside, Farmland, Parking, Town, Mountain, and Suburb.

Seven scenes, with Ground Truth, BlockGaussian, MonoGS*, VINGS-Mono, and ProAGS in that order
From left to right: Ground Truth, BlockGaussian, MonoGS*, VINGS-Mono, and ProAGS. The ProAGS images are rendered without the optional global refinement.

Efficiency

Only the active region is loaded onto the GPU. The full Gaussian map grows in CPU memory as the UAV observes more of the scene.

Measured on one NVIDIA RTX 3090 (24 GB), Intel i9-14900KF, and 128 GB RAM.

Original runtime and memory results for Factory, Countryside, Farmland, Parking, Town, Mountain, and Suburb

Average processing time ranges from 0.787 to 0.855 seconds per input frame. With photograph intervals of approximately 0.7–1 second, this supports near-online progressive mapping. These are processing times, not rendering frame rates, and assume camera poses are already available.

Memory and computation on Scene Suburb

Forward processing time in milliseconds over input framesBackward processing time in milliseconds over input framesGPU memory in gigabytes over input framesCPU memory in gigabytes over input framesActive local Gaussian count over input framesGlobal Gaussian count over input frames
Horizontal axes: input frame index. Forward/backward times are in milliseconds; memory is in GB. Local Gaussian count and VRAM fluctuate as cells are loaded and unloaded, while CPU memory grows with the global map.

How well are previous views retained?

On Suburb, each keyframe is tracked for the next 200 frames after it first enters optimization. Metrics are averaged over the tracked keyframes. Weighted replay better preserves previously observed views throughout this window.

Suburb forgetting curves: weighted replay preserves PSNR and SSIM and keeps LPIPS lower than random replay or no replay
PSNR and SSIM: higher is better. LPIPS: lower is better. Horizontal axes count subsequent frames.

Limitations

Where reliable geometry remains difficult

ProAGS targets progressive mapping from accurately posed aerial photographs. Pose estimation is outside this pipeline, and the reported throughput does not establish strict real-time operation.

Weak texture and reflections

Water surfaces make multi-view matching unreliable. Confidence filtering can then leave too few valid depth samples to initialize Gaussians, producing holes, blur, or floating artifacts.

Terrain and memory

Extreme terrain relief or large vertical structures can challenge the horizontal BEV partition and local visibility estimate. The global map still consumes increasing CPU memory as the scene expands.

Two water-dominated failure cases: ground truth, MVS depth, valid depth after confidence filtering, and ProAGS render
Lakes and rivers expose the dependence on reliable multi-view depth. White regions in the valid-depth column indicate missing support after filtering.

Citation

@article{wu2026progressive,
  title={Progressive Gaussian Splatting for High-Fidelity 3D Reconstruction of Large-Scale Remote Sensing Scenes: Practical, scalable, and compatible with standard aerial imaging workflows},
  author={Wu, Yongchang and Shi, Zhenwei and Zou, Zhengxia},
  journal={IEEE Geoscience and Remote Sensing Magazine},
  volume={14},
  number={4},
  pages={240--258},
  year={2026},
  publisher={IEEE}
}

Figure

Open image