吉林大学学报(地球科学版) ›› 2025, Vol. 55 ›› Issue (5): 1629-1643.doi: 10.13278/j.cnki.jjuese.20230139

• 地质工程与环境工程 • 上一篇    下一篇

基于机器学习的富硒土壤预测模型的构建与比较——以江西省信丰县油山地区为例

杨兰1,王运1,邹勇军2,胡宝群1,李满根1,张安1,朱满怀1   

  1. 1.东华理工大学江西省数字国土重点实验室,南昌330013

    2.江西省地质勘查院地质环境监测所,南昌330001

  • 出版日期:2025-09-26 发布日期:2025-11-15
  • 基金资助:
    中国地质调查局项目(DD20160321);江西省重点研发计划项目(20203BBG72W011);东华理工大学博士启动基金项目(DHBK2019051);东华理工大学江西省数字国土重点实验室开放研究基金资助项目(DLLJ202205);江西省研究生创新专项资金项目(YC2022-S600)

Construction and Comparison of Models for Predicting Selenium Rich Soil Based on Machine Learning: A Case Study of Youshan Area,Xinfeng County, Jiangxi Province

Yang Lan1,Wang Yun1,Zou Yongjun2,Hu Baoqun1,Li Mangen1,Zhang An1,Zhu Manhuai1   

  1. 1. Key Laboratory of Digital Land and Resources of Jiangxi Province,East China University of Technology,

    Nanchang 330013,China

    2. Geological Environment Monitoring Institute of Jiangxi Geological Exploration Institute,Nanchang 330001,China

  • Online:2025-09-26 Published:2025-11-15
  • Supported by:
    Supported by the Project of China Geological Survey (DD20160321),the Key Research and Development Plan Project of Jiangxi Province (20203BBG72W011),the Doctoral Startup Fund of East China University of Technology (DHBK2019051), the Project of Key Laboratory for Digital Land and Resources of Jiangxi Province, East China University of Technology (DLLJ202205) and the Project of Jiangxi Postgraduate Innovation Special Fund (YC2022-S600)

摘要:

利用未知硒数据快速、高效、精准地圈定富硒土壤,需构建预测富硒土壤的最佳模型。从1 277个1∶5万表层土壤的地球化学数据中选取502个数据组成数据集,以w(Zn)、w(K2O)、w(P)、w(Mo)、w(Mn)、w(Cr)、pH、D(泥盆系)为自变量,以是否富Se为因变量,运用SPSS Modeler 18软件构建二元Logistic回归模型、多层感知器神经网络模型、随机森林模型及支持向量机模型(包括线性、多项式、径向基和Sigmoid核函数),并通过35组土壤样品实测数据进行验证。结果表明:二元Logistic回归模型、多层感知器神经网络模型、随机森林模型及(线性、多项式、径向基、Sigmoid)支持向量机模型的预测准确率和验证总体准确率分别为88.8%和94.3%、91.0%和97.1%、96.6%和97.1%、87.9%和97.1%、86.1%和94.3%、86.9%和94.3%、80.3%和91.4%;以上模型的曲线下面积(AUC)值分别为0.948、0.950、0.993、0.937、0.945、0.928和0.873,随机森林模型的准确率和稳定性最佳。同时,本次研究发现了清洁富硒土壤及绿色富硒山稻,表明该方法在富硒土壤预测中具有可行性,且可进一步拓展到地质找矿及环境监测等领域。

关键词: 富硒土壤, 机器学习, 二元Logistic回归模型, 多层感知器神经网络模型, 随机森林模型, 支持向量机模型

Abstract:

In order to find selenium rich soil quickly, efficiently and accurately using selenium free data, it is necessary to build the best model to predict selenium rich soil. 502 data sets were selected from 1 277 1∶50 000 surface soil geochemical data. With w(Zn),w(K2O),w(P),w(Mo),w(Mn),w(Cr),pH,D(Devonian) as independent variables and Se rich or not as dependent variables, SPSS Modeler 18 software was used to build binary Logistic regression model, multi-layer perceptron neural network model, random forest model and support vector machine model (linear, multinomial, radial basis function, Sigmoid) for predicting Se rich soil, and the measured data of 35 soil samples were used for verification. The results show that, using binary Logistic regression model, multilayer perceptron neural network model, random forest model and support vector machine model (linear, polynomial, radial basis function, Sigmoid), the overall accuracy of prediction and verification of the seven prediction models and were 88.8% and 94.3%, 91.0% and 97.1%, 96.6% and 97.1%, 87.9% and 97.1%, 86.1% and 94.3%, 86.9% and 94.3%, 80.3% and 91.4%. The AUC were 0.948, 0.950, 0.993, 0.937, 0.945, 0.928 and 0.873, respectively. The accuracy and stability of the random forest model are the best. Meanwhile, this study identified clean selenium-rich soil and green selenium-rich mountain rice, indicating that this method is feasible in the prediction of selenium-rich soil, and it can be further extended to geological prospecting and environmental monitoring.

Key words: selenium rich soil, machine learning, binary Logistic regression model, multilayer perceptron neural network model, random forest, support vector machine model

中图分类号: 

  • P59
[1] 张恩威, 孟庆涛, 唐佰强, 胡菲, 党微.  基于机器学习的TOC测井预测方法:以松辽盆地南部青山口组一段为例[J]. 吉林大学学报(地球科学版), 2026, 56(3): 1062-1075.
[2] 王菲, 龙欣雨, 唐杰, 郭鹏, 许文良. 古亚洲洋东段的剪刀式闭合历史——来自华北克拉通北缘东段及邻区三叠纪地壳厚度空间变异的制约[J]. 吉林大学学报(地球科学版), 2026, 56(1): 101-117.
[3] 韩复兴, 刘水源, 高正辉, 韩江涛, 张涛, 尚浩. 基于机器学习的微动HVSR数据干扰信号压制方法[J]. 吉林大学学报(地球科学版), 2025, 55(6): 2153-2163.
[4] 曹志民, 张丽, 郑兵, 韩建. 基于SMOTE平衡数据的极端随机树岩性识别[J]. 吉林大学学报(地球科学版), 2025, 55(4): 1372-1386.
[5] 曹志民, 丁璐, 韩建, 郝乐川, .

基于集成机器学习的测井曲线大尺度差异超分辨 [J]. 吉林大学学报(地球科学版), 2025, 55(2): 670-685.

[6] 吕华星, 陈兆明, 张振波, 姜大朋, 李克成, 郭伟. 机器学习高分辨融合反演在地层对比中的应用——以珠江口盆地开平凹陷开平A构造带为例[J]. 吉林大学学报(地球科学版), 2025, 55(1): 289-297.
[7] 王新领, 祝新益, 张宏兵, 孙博, 许可欣.

基于随机树嵌入的随钻测井岩性识别方法 [J]. 吉林大学学报(地球科学版), 2024, 54(2): 701-708.

[8] 杨国华, 李婉露, 孟博. 基于机器学习方法的地下水氨氮时空分布规律[J]. 吉林大学学报(地球科学版), 2022, 52(6): 1982-1995.
[9] 侯贤沐, 王付勇, 宰芸, 廉培庆. 基于机器学习和测井数据的碳酸盐岩孔隙度与渗透率预测[J]. 吉林大学学报(地球科学版), 2022, 52(2): 644-653.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
[1] 张世广, 柳成志, 卢双舫, 张雁, 吴高平, 刘秋宏. 高分辨率层序地层学在河、湖、三角洲复合沉积体系的应用--以朝阳沟油田扶余油层开发区块为例[J]. J4, 2009, 39(3): 361 -368 .
[2] 李碧乐,沈鑫,陈广俊,杨延乾,李永胜. 青海东昆仑阿斯哈金矿Ⅰ号脉成矿流体地球化学特征和矿床成因[J]. 吉林大学学报(地球科学版), 2012, 42(6): 1676 -1687 .
[3] 杨晓平,李仰春,柳 震, 汪 岩,王洪杰. 黑龙江东部鸡西盆地构造层序划分与盆地动力学演化[J]. J4, 2005, 35(05): 616 -621 .
[4] 李发文, 冯平, 张超. 天津北三河地区垂向耦合产流模型及应用[J]. J4, 2011, 41(2): 459 -464 .
[5] 王立军,刘国才,黄继国,李跃迁,丛颖,沈照理,赵晓波. 接触氧化技术在公园景观水体功能恢复中的运用试验研究[J]. J4, 2006, 36(03): 458 -461 .
[6] 雷如雄,吴昌志,屈迅,顾连兴,陈刚,吾尔娜,孙洪涛,刘国宁. 中天山天湖东铁钼矿含矿片麻状花岗岩年代学、地球化学和锆石Hf同位素-对于中天山早古生代构造演化的启示[J]. 吉林大学学报(地球科学版), 2014, 44(5): 1540 -1552 .
[7] 初凤友,胡大千,姚杰. 中太平洋YJB海山富钴结核矿物组成与元素地球化学[J]. J4, 2007, 37(1): 8 -0014 .
[8] 周燕,郑培玺,王铁夫,张延洁. 招平断裂带上盘金矿床氢氧同位素地质特征[J]. J4, 2007, 37(4): 668 -0671 .
[9] 唐华风,王璞珺,姜传金,刘杰,张庆晨,冯有良. 松辽盆地火山岩相地震特征及其与控陷断裂的关系[J]. J4, 2007, 37(1): 73 -0078 .
[10] 范晓敏,李舟波. 裂缝性碳酸盐岩储层声波时差曲线的波动和增幅分析[J]. J4, 2007, 37(1): 168 -0173 .