Large-scale scene reconstruction · 3D Gaussian Splatting

BlockGaussian: Efficient Large-Scale Scene Novel View Synthesis
via Adaptive Block-Based Gaussian Splatting

Yongchang WuZipeng QiZhenwei ShiZhengxia Zou*

Beihang University

* Corresponding author

A reconstructed city with novel views from multiple camera positions

Large-scale novel view synthesis with limited computing resources. Balanced blocks, focused optimization, and consistent fusion.

Single-GPU training

Train blocks sequentially on a consumer GPU, or in parallel across multiple GPUs.

3.8–6.6× faster

Reconstruction vs. VastGaussian across six scenes on 8× RTX 5090. Table III →

73.0–79.8% retained

Optimized Gaussians retained after boundary pruning across four scenes. Table V.

Abstract

Recent advances in 3D Gaussian Splatting have demonstrated remarkable potential in novel view synthesis. While the divide-and-conquer paradigm has enabled large-scale scene reconstruction, challenges remain in scene partitioning, optimization, and merging under GPU memory constraints. We introduce BlockGaussian, a framework for large-scale novel view synthesis that enables efficient scene reconstruction on a single consumer-grade GPU. Specifically, we propose a content-aware scene partitioning strategy that considers scene content distribution and computational load to balance reconstruction workloads across blocks and improve computational efficiency. Furthermore, a visibility-aware block optimization strategy alleviates the mismatch between block-local representations and image supervision, thereby reducing redundant computation. To preserve rendering quality after block fusion, a pseudo-view geometry constraint suppresses free-space floaters through local reprojection regularization without requiring additional captured views. Extensive experiments demonstrate that our approach substantially reduces reconstruction time while achieving rendering quality competitive with state-of-the-art methods. By decomposing large-scale reconstruction into localized block-wise subproblems, our framework reduces the per-block optimization burden and enables efficient parallel reconstruction across multiple GPUs.

Balanced blocks. Efficient reconstruction.

Splitting a scene is only the beginning. Efficient reconstruction also needs balanced workloads, appropriate supervision, and reliable geometry at fusion.

01

Uneven workloads

Uniform spatial grids can assign very different reconstruction loads to each block. Parallel training then waits for the slowest block.

02

Incomplete scene context

A block covers only part of a training image. Fitting the entire image with block-local Gaussians encourages redundant out-of-block growth.

03

Unstable free space

Floaters may fit training views but become visible from novel viewpoints after independently trained blocks are merged.

BlockGaussian pipeline: content-aware partitioning, visibility-aware optimization and scene fusion

01 Partition by content

Use sparse SfM points as a proxy for scene complexity. Recursively split dense regions and assign views through visibility, without a global coarse-Gaussian optimization stage.

02 Optimize useful content

Auxiliary Gaussians represent visible out-of-block regions. Only block Gaussians are densified; auxiliary Gaussians remain compact and are removed before merging.

03 Regularize free space

Perturb existing training views and warp the rendered pseudo views back to their reference views. Local reprojection consistency discourages unstable free-space Gaussians.

Explore content-aware partitioning

The partition adapts to the distribution of sparse scene content. Building, Rubble, Residence, and MatrixCity-Street each use 8 blocks; Sci-Art uses 4 and MatrixCity-Aerial uses 16. SfM point quality remains important for estimating reconstruction complexity.

Content-adaptive partitions for Building, Rubble, Sci-Art, Residence, MatrixCity-Street and MatrixCity-Aerial
Partition layouts across aerial and street scenes.
How pseudo-view regularization works

A small camera-pose perturbation creates a local pseudo view from an existing training view. Its rendered depth reprojects the image back to the reference view, where consistency with the ground truth constrains unstable free-space Gaussians without requiring additional captured views.

Local camera perturbation and depth-based reprojection for the pseudo-view geometry constraint
Pseudo-view geometry constraint.

Rendering quality

From aerial scenes to street views.

Evaluate quality alongside speed. BlockGaussian preserves detailed urban structure while achieving competitive results across real aerial datasets and synthetic city scenes.

Table I: rendering quality on Mill19 and UrbanScene3D
Qualitative comparison across Mill19 and UrbanScene3D against Mega-NeRF, 3DGS, CityGaussian and DOGS
Real aerial scenes. Detailed crops show structural boundaries, vegetation and facade detail.

Efficiency

Less redundant work. Faster reconstruction.

BlockGaussian combines lightweight partitioning, localized optimization, and balanced block workloads. The gain is visible in both reconstruction time and the proportion of optimized Gaussians that survive merging.

Table III: reconstruction time, Gaussian count and model storage across six scenes

Across six scenes, reconstruction is 3.8–6.6× faster than VastGaussian on the shared 8-GPU platform. Speedups use the times in Table III.

Balanced workloads across blocks.

Per-block optimization time for four scenes

Training behavior of block and auxiliary Gaussians.

Auxiliary Gaussians show larger fluctuations in position and log-scale gradients, while block Gaussians become more stable during training. This supports treating out-of-block context separately from the block representation.

Position and log-scale gradient norms during training
Gradient behavior for one block in MatrixCity-Aerial. Blue: block Gaussians; orange: auxiliary Gaussians.

Complementary components, measured step by step.

The Rubble ablation follows the adopted configuration sequence. Auxiliary context improves efficiency, while the scheduled pseudo-view loss trades some computation for better rendering quality and fewer free-space artifacts.

Table VI: complete algorithm-component ablation on Rubble
Comparison protocol

Runtime measurements use 8× NVIDIA RTX 5090 GPUs. Reproduced Gaussian methods are trained for 40,000 iterations. Mill19 and UrbanScene3D images use 4× downsampling; reproduced DOGS also uses 4× rather than its original 6× setting. VastGaussian uses an unofficial implementation; CityGaussian, CityGaussianV2 and DOGS use official codebases. Scenes are re-partitioned for the efficiency comparison. Our default batch size is 1.

Explore the reconstructions

See the scene. Inspect the details.

Building · BlockGaussian novel-view rendering. Use the player controls to start or pause.

Interactive image comparison

3DGS BlockGaussian

Scene Building: left is 3DGS, right is BlockGaussian. Drag the divider to compare.

Dynamic comparison

Left: VastGaussian · Right: BlockGaussian

Scope and limitations

What the framework targets.

BlockGaussian targets reconstruction efficiency. Final Gaussian counts, storage, rendering memory, and full-scene rendering latency can still grow with scene scale. Rendering-oriented techniques such as level of detail and compression are complementary directions.

Rendering quality does not guarantee geometric accuracy. The framework also relies on reliable scene initialization, so challenging image conditions or incomplete scene estimates can limit reconstruction.

BibTeX

@article{wu2025blockgaussian,
  title={Blockgaussian: Efficient large-scale scene novel view synthesis via adaptive block-based gaussian splatting},
  author={Wu, Yongchang and Qi, Zipeng and Shi, Zhenwei and Zou, Zhengxia},
  journal={arXiv preprint arXiv:2504.09048},
  year={2025}
}

Figure

Open image