Journal of Jilin University(Engineering and Technology Edition) ›› 2026, Vol. 56 ›› Issue (7): 2020-2025.doi: 10.13229/j.cnki.jdxbgxb.20241303

Previous Articles    

Real-time scene reconstruction algorithm based on neural voxel fusion splatting learning

Peng-kun WU1(),Xing-chen WU2()   

  1. 1.College of Design and Art,Shenyang Aerospace University Shenyang,Shenyang 110136,China
    2.JiHua Laboratory,Foshan 528200,China
  • Received:2024-12-04 Online:2026-07-01 Published:2026-08-12
  • Contact: Xing-chen WU E-mail:2677675362@qq.com;wuxingchen3687@sina.com

Abstract:

Three-dimensional scene reconstruction from a set of two-dimensional images remains a significant and challenging research task in computer vision. Existing algorithms struggle to effectively model the relationships between images captured at different times, resulting in poor rendering quality for large-scale environments. To address this challenge, this paper proposes a neural surfels fusion splatting algorithm for real-time scene reconstruction. The algorithm progressively builds and updates scene models from streaming input data, incorporating neural surfel representations based on 2D Gaussian primitives and a fusion mechanism based on Gated Recurrent Units (GRU). First, the algorithm achieves perspective-correct rendering by integrating Gaussian surfel splatting into the rendering pipeline and employing explicit ray-splat intersections, thereby enhancing depth consistency and geometric fidelity across multiple viewpoints. Then, to further improve the quality of reconstructed surfaces, the method combines depth pulling and normal consistency regularization, promoting smoother surface reconstruction and facilitating high-quality mesh extraction. The algorithm's performance was evaluated on two large-scale indoor/outdoor scene reconstruction datasets: DeepBlending and Tanks&Temples. Extensive experimental results demonstrate that the algorithm outperforms existing methods in both quantitative and qualitative assessments.

Key words: pose estimation, feature alignment, sequence information modeling, CNN

CLC Number: 

  • TP391.4

Fig.1

Overall pipeline of real-time real-world scene reconstruction via neural facet fusion splatting"

Table 1

Quantitative results comparison between our algorithm and existing methods on the deepblending dataset"

方法PlayroomDr. Johnson
PSNRSSIMLPIPSPSNRSSIMLPIPS
NeuS1427.690.8830.28426.300.8640.314
3DG730.010.8980.26028.450.8880.277
SuGaR830.120.8970.25728.710.8880.268
本文算法30.870.8980.25329.150.8890.262

Fig.2

Quantitative results comparison between this algorithm and existing methods on the deepblending dataset"

Table 2

Quantitative results comparison between this method and existing approaches on the tanks & temples dataset"

方法TrainTruck
PSNRSSIMLPIPSPSNRSSIMLPIPS
NeuS1418.130.6940.33520.080.7870.235
3DG719.860.7490.27422.240.8220.187
SuGaR820.500.7630.25822.670.8270.174
本文算法21.360.7650.25623.510.8310.169

Fig.3

Visualization results of the algorithm on outdoor scenes in the Tanks & Temples dataset"

Table 3

Ablation experiment"

方 法PSNRSSIMLPIPS
使用神经辐射场表示29.830.8920.257
删除基于门控循环单元的融合机制30.120.8850.261
完整模型30.870.8980.253
[1] Cheng J, Zhang L Y, Chen Q H, et al. A review of visual SLAM methods for autonomous driving vehicles[J]. Engineering Applications of Artificial Intelligence, 2022, 114: 104992.
[2] Kazerouni I A, Fitzgerald L, Dooly G, et al. A survey of state-of-the-art on visual SLAM[J]. Expert Systems with Applications, 2022, 205: 117734.
[3] Lee T J, Kim C H, Cho D D. A monocular vision sensor-based efficient SLAM method for indoor service robots[J]. IEEE Transactions on Industrial Electronics, 2018, 66(1): 318-328.
[4] Wei Y, Liu S H, Rao Y M, et al. Nerfingmvs: Guided optimization of neural radiance fields for indoor multi-view stereo[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021: 5610-5619.
[5] Chen Z, Wang C, Guo Y C, et al. Structnerf: Neural radiance fields for indoor scenes with structural hints[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(12): 15694-15705.
[6] Liu J L, Nie Q, Liu Y, et al. Nerf-loc: Visual localization with conditional neural radiance field[C]∥IEEE International Conference on Robotics and Automation(ICRA), London, United Kingdom, 2023: 9385-9392.
[7] Kerbl B, Kopanas G, Leimkühler T, et al. 3D gaussian splatting for real-time radiance field rendering[J]. ACM Trans. Graph., 2023, 42(4): 139:1-14.
[8] Stier N, Ranjan A, Colburn A, et al. Finerecon: Depth-aware feed-forward network for detailed 3d reconstruction[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2023: 18377-18386.
[9] Wei W F, Wang J, Xie X, et al. Real-time dense visual slam with neural factor representation[J]. Electronics, 2024, 13(16): 3332.
[11] Tretschk E, Kairanda N, Br M, et al. State of the art in dense monocular non‐rigid 3D reconstruction[J]. Computer Graphics Forum, 2023, 42(2): 485-520.
[12] Hedman P, Philip J, Price T, et al. Deep blending for free-viewpoint image-based rendering[J]. ACM Transactions on Graphics, 2018, 37(6): 1-15.
[13] Knapitsch A, Park J, Zhou Q Y, et al. Tanks and temples: benchmarking large-scale scene reconstruction[J]. ACM Transactions on Graphics, 2017, 36(4): 1-13.
Guédon A, Lepetit V. Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 2024: 5354-5363.
[14] Wang P, Liu L, Liu Y, et al. Neus: learning neural implicit surfaces by volume rendering for multi-view reconstruction[J]. arXiv preprint arXiv:, 2021.
[15] Zhang R, Isola P, Efros A A, et al. The unreasonable effectiveness of deep features as a perceptual metric[C]∥Proceedings of the IEEE conference on computer vision and pattern recognition, Salt Lake City, UT, USA, 2018: 586-595.
[1] Feng SHI,Peng NIU,Min FAN. Uneven deformation detection of highway subgrade and pavement based on Faster R-CNN algorithm [J]. Journal of Jilin University(Engineering and Technology Edition), 2026, 56(7): 1950-1957.
[2] Jia-bao LI,Cheng-jun WANG,Wen-hang SU. 2D human pose estimation algorithm based on adaptive parameterized non-maximum suppression [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(7): 2425-2433.
[3] Xue-jun LI,Lin-fei QUAN,Dong-mei LIU,Shu-you YU. Improved Faster⁃RCNN algorithm for traffic sign detection [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(3): 938-946.
[4] Hua CAI,Rui-kun ZHU,Qiang FU,Wei-gang WANG,Zhi-yong MA,Jun-xi SUN. Human pose estimation corrector algorithm based on implicit key point interconnection [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(3): 1061-1071.
[5] Xin GUAN,Zi-jian ZHOU,Qiang LI. Human pose estimation based on graph structure guidance and location information enhancement [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(10): 3283-3295.
[6] Wei WANG,Yu-jie SUN,Xin WANG. Lightweight frequency and spatial feature fused multi-scale remote sensing scene classification network [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(10): 3361-3371.
[7] Yu WANG,Kai ZHAO. Postprocessing of human pose heatmap based on sub⁃pixel location [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(5): 1385-1392.
[8] Ren-xiang CHEN,Chao-chao HU,Xiao-lin HU,Li-xia YANG,Jun ZHANG,Jia-le HE. Driver distracted driving detection based on improved YOLOv5 [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(4): 959-968.
[9] Fang-shi WANG,Peng BAO. Intelligent Recognition of sensitive small targets with fine grains in complex background remote sensing images [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(11): 3289-3295.
[10] Yue-lin CHEN,Zhu-cheng GAO,Xiao-dong CAI. Long text semantic matching model based on BERT and dense composite network [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(1): 232-239.
[11] Lian-ming WANG,Xin WU. Method for 3D motion parameter measurement based on pose estimation [J]. Journal of Jilin University(Engineering and Technology Edition), 2023, 53(7): 2099-2108.
[12] Ming-hua GAO,Can YANG. Traffic target detection method based on improved convolution neural network [J]. Journal of Jilin University(Engineering and Technology Edition), 2022, 52(6): 1353-1361.
[13] Xian-tong LI,Wei QUAN,Hua WANG,Peng-cheng SUN,Peng-jin AN,Yong-xing MAN. Route travel time prediction on deep learning model through spatiotemporal features [J]. Journal of Jilin University(Engineering and Technology Edition), 2022, 52(3): 557-563.
[14] You QU,Wen-hui LI. Single-stage rotated object detection network based on anchor transformation [J]. Journal of Jilin University(Engineering and Technology Edition), 2022, 52(1): 162-173.
[15] Hou⁃jie LI,Fa⁃sheng WANG,Jian⁃jun HE,Yu ZHOU,Wei LI,Yu⁃xuan DOU. Pseudo sample regularization Faster R⁃CNN for traffic sign detection [J]. Journal of Jilin University(Engineering and Technology Edition), 2021, 51(4): 1251-1260.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!