吉林大学学报(工学版) ›› 2026, Vol. 56 ›› Issue (7): 2020-2025.doi: 10.13229/j.cnki.jdxbgxb.20241303

• 计算机科学与技术 • 上一篇    

基于神经面元融合溅射学习的实时真实场景重建算法

邬鹏坤1(),吴星辰2()   

  1. 1.沈阳航空航天大学 设计艺术学院,沈阳 110136
    2.季华实验室,广东 佛山 528200
  • 收稿日期:2024-12-04 出版日期:2026-07-01 发布日期:2026-08-12
  • 通讯作者: 吴星辰 E-mail:2677675362@qq.com;wuxingchen3687@sina.com
  • 作者简介:邬鹏坤(1980-),男,副教授. 研究方向:数字媒体艺术.E-mail: 2677675362@qq.com
  • 基金资助:
    辽宁省教育厅一般科研项目(LJKMR20220546)

Real-time scene reconstruction algorithm based on neural voxel fusion splatting learning

Peng-kun WU1(),Xing-chen WU2()   

  1. 1.College of Design and Art,Shenyang Aerospace University Shenyang,Shenyang 110136,China
    2.JiHua Laboratory,Foshan 528200,China
  • Received:2024-12-04 Online:2026-07-01 Published:2026-08-12
  • Contact: Xing-chen WU E-mail:2677675362@qq.com;wuxingchen3687@sina.com

摘要:

基于一组二维图像对三维场景进行重建的过程中,现有算法难以建模不同时间图片间的关联关系,导致产生较差的大型环境渲染结果,针对该问题,提出一种用于实时真实场景重建的神经面元融合溅射算法。该算法从流式输入数据中逐步构建和更新场景模型,包括基于2D高斯基元的神经面元表示和基于门控循环单元(GRU)的融合机制。首先,通过将高斯面元溅射集成到渲染管道中并采用显式光线溅射交点,实现了透视校正渲染,从而增强了跨多个视点的深度一致性和几何保真度;然后,为了进一步提高重建表面的质量,该方法结合了深度拉取和法线一致性正则化,从而促进了更平滑的表面重建并有助于提取高质量的网格;最后,在DeepBlending和Tanks&Temples两个大型室内/室外场景重建数据集对算法性能进行测试。实验结果表明:该算法在定量和定性评估方面均优于现有算法。

关键词: 姿态估计, 特征对齐, 序列信息建模, 卷积神经网络

Abstract:

Three-dimensional scene reconstruction from a set of two-dimensional images remains a significant and challenging research task in computer vision. Existing algorithms struggle to effectively model the relationships between images captured at different times, resulting in poor rendering quality for large-scale environments. To address this challenge, this paper proposes a neural surfels fusion splatting algorithm for real-time scene reconstruction. The algorithm progressively builds and updates scene models from streaming input data, incorporating neural surfel representations based on 2D Gaussian primitives and a fusion mechanism based on Gated Recurrent Units (GRU). First, the algorithm achieves perspective-correct rendering by integrating Gaussian surfel splatting into the rendering pipeline and employing explicit ray-splat intersections, thereby enhancing depth consistency and geometric fidelity across multiple viewpoints. Then, to further improve the quality of reconstructed surfaces, the method combines depth pulling and normal consistency regularization, promoting smoother surface reconstruction and facilitating high-quality mesh extraction. The algorithm's performance was evaluated on two large-scale indoor/outdoor scene reconstruction datasets: DeepBlending and Tanks&Temples. Extensive experimental results demonstrate that the algorithm outperforms existing methods in both quantitative and qualitative assessments.

Key words: pose estimation, feature alignment, sequence information modeling, CNN

中图分类号: 

  • TP391.4

图1

基于神经面元融合溅射学习的实时真实场景重建算法整体流程图"

表1

本文算法与现有方法在DeepBlending数据集的定量结果比较"

方法PlayroomDr. Johnson
PSNRSSIMLPIPSPSNRSSIMLPIPS
NeuS1427.690.8830.28426.300.8640.314
3DG730.010.8980.26028.450.8880.277
SuGaR830.120.8970.25728.710.8880.268
本文算法30.870.8980.25329.150.8890.262

图2

本文算法在DeepBlending数据集上室内场景的可视化结果"

表2

本文算法与现有方法在Tanks&Temples数据集的定量结果比较"

方法TrainTruck
PSNRSSIMLPIPSPSNRSSIMLPIPS
NeuS1418.130.6940.33520.080.7870.235
3DG719.860.7490.27422.240.8220.187
SuGaR820.500.7630.25822.670.8270.174
本文算法21.360.7650.25623.510.8310.169

图3

本文算法在Tanks&Temples数据集上室外场景的可视化结果"

表3

消融实验"

方 法PSNRSSIMLPIPS
使用神经辐射场表示29.830.8920.257
删除基于门控循环单元的融合机制30.120.8850.261
完整模型30.870.8980.253
[1] Cheng J, Zhang L Y, Chen Q H, et al. A review of visual SLAM methods for autonomous driving vehicles[J]. Engineering Applications of Artificial Intelligence, 2022, 114: 104992.
[2] Kazerouni I A, Fitzgerald L, Dooly G, et al. A survey of state-of-the-art on visual SLAM[J]. Expert Systems with Applications, 2022, 205: 117734.
[3] Lee T J, Kim C H, Cho D D. A monocular vision sensor-based efficient SLAM method for indoor service robots[J]. IEEE Transactions on Industrial Electronics, 2018, 66(1): 318-328.
[4] Wei Y, Liu S H, Rao Y M, et al. Nerfingmvs: Guided optimization of neural radiance fields for indoor multi-view stereo[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021: 5610-5619.
[5] Chen Z, Wang C, Guo Y C, et al. Structnerf: Neural radiance fields for indoor scenes with structural hints[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(12): 15694-15705.
[6] Liu J L, Nie Q, Liu Y, et al. Nerf-loc: Visual localization with conditional neural radiance field[C]∥IEEE International Conference on Robotics and Automation(ICRA), London, United Kingdom, 2023: 9385-9392.
[7] Kerbl B, Kopanas G, Leimkühler T, et al. 3D gaussian splatting for real-time radiance field rendering[J]. ACM Trans. Graph., 2023, 42(4): 139:1-14.
[8] Stier N, Ranjan A, Colburn A, et al. Finerecon: Depth-aware feed-forward network for detailed 3d reconstruction[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2023: 18377-18386.
[9] Wei W F, Wang J, Xie X, et al. Real-time dense visual slam with neural factor representation[J]. Electronics, 2024, 13(16): 3332.
[11] Tretschk E, Kairanda N, Br M, et al. State of the art in dense monocular non‐rigid 3D reconstruction[J]. Computer Graphics Forum, 2023, 42(2): 485-520.
[12] Hedman P, Philip J, Price T, et al. Deep blending for free-viewpoint image-based rendering[J]. ACM Transactions on Graphics, 2018, 37(6): 1-15.
[13] Knapitsch A, Park J, Zhou Q Y, et al. Tanks and temples: benchmarking large-scale scene reconstruction[J]. ACM Transactions on Graphics, 2017, 36(4): 1-13.
Guédon A, Lepetit V. Sugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 2024: 5354-5363.
[14] Wang P, Liu L, Liu Y, et al. Neus: learning neural implicit surfaces by volume rendering for multi-view reconstruction[J]. arXiv preprint arXiv:, 2021.
[15] Zhang R, Isola P, Efros A A, et al. The unreasonable effectiveness of deep features as a perceptual metric[C]∥Proceedings of the IEEE conference on computer vision and pattern recognition, Salt Lake City, UT, USA, 2018: 586-595.
[1] 侯越,张鑫,武月. 基于时空动态约束图反馈的交通流预测[J]. 吉林大学学报(工学版), 2026, 56(1): 183-198.
[2] 田婧,马社强,宋现敏,赵丹,陈发城. 稀疏数据下的交通状态估计自适应卷积网络[J]. 吉林大学学报(工学版), 2025, 55(8): 2579-2587.
[3] 肖红,刘显德. 基于混合智能的学习状态实时采集与动态分析方法[J]. 吉林大学学报(工学版), 2025, 55(7): 2402-2408.
[4] 李家宝,王成军,苏文杭. 基于自适应参数化非极大值抑制的二维人体姿态估计算法[J]. 吉林大学学报(工学版), 2025, 55(7): 2425-2433.
[5] 冯志刚,王首起,于明月. 基于变分模态提取及轻量级网络的滚动轴承故障诊断[J]. 吉林大学学报(工学版), 2025, 55(6): 1883-1891.
[6] 杜睿山,王紫珊. 基于时空注意力的多视角人脸表情识别算法[J]. 吉林大学学报(工学版), 2025, 55(6): 2097-2102.
[7] 才华,朱瑞昆,付强,王伟刚,马智勇,孙俊喜. 基于隐式关键点互联的人体姿态估计矫正器算法[J]. 吉林大学学报(工学版), 2025, 55(3): 1061-1071.
[8] 罗维薇,刘长龙,雷琴. 基于预处理层增强和注意力机制的空域图像隐写分析[J]. 吉林大学学报(工学版), 2025, 55(12): 4024-4033.
[9] 周求湛,牟岩,武慧南,陈霄,汪锋,李琛,张雯,刘萍萍,王聪. 基于RBVS和CBCNN的风机叶片故障检测和分类方法[J]. 吉林大学学报(工学版), 2025, 55(10): 3119-3130.
[10] 关欣,周子健,李锵. 基于图结构引导和位置信息强化的人体姿态估计[J]. 吉林大学学报(工学版), 2025, 55(10): 3283-3295.
[11] 王威,孙钰洁,王新. 频率和空间特征融合的轻量级多尺度遥感图像场景分类网络[J]. 吉林大学学报(工学版), 2025, 55(10): 3361-3371.
[12] 汪豪,赵彬,刘国华. 基于时间和运动增强的视频动作识别[J]. 吉林大学学报(工学版), 2025, 55(1): 339-346.
[13] 胡宏宇,张争光,曲优,蔡沐雨,高菲,高镇海. 基于双分支和可变形卷积网络的驾驶员行为识别方法[J]. 吉林大学学报(工学版), 2025, 55(1): 93-104.
[14] 特木尔朝鲁朝鲁,张亚萍. 基于卷积神经网络的无线传感器网络链路异常检测算法[J]. 吉林大学学报(工学版), 2024, 54(8): 2295-2300.
[15] 赵宏伟,武鸿,马克,李海. 基于知识蒸馏的图像分类框架[J]. 吉林大学学报(工学版), 2024, 54(8): 2307-2312.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!