吉林大学学报(医学版) ›› 2026, Vol. 52 ›› Issue (4): 1086-1095.doi: 10.13481/j.1671-587X.20260421

• 临床研究 • 上一篇    

基于临床实验室常规检测数据的食管癌机器学习预测模型的构建和验证

李蕊,郝晓燕,杨柳,刘家云,何睦()   

  1. 空军军医大学第一附属医院检验科,陕西 西安 710032
  • 收稿日期:2025-11-06 接受日期:2026-01-04 出版日期:2026-07-28 发布日期:2026-07-27
  • 通讯作者: 何睦 E-mail:394092811@qq.com
  • 作者简介:李 蕊(1988-),女,内蒙古自治区赤峰市人,主管技师,医学硕士,主要从事临床检验诊断学方面的研究。
  • 基金资助:
    国家重点研发计划项目(2022YFC3602301)

Construction and validation of machine learning-based prediction models for esophageal cancer using routine clinical laboratory data

Rui LI,Xiaoyan HAO,Liu YANG,Jiayun LIU,Mu HE()   

  1. Department of Clinical Laboratory Medicine,First Affiliated Hospital,Air Force Medical University,Xi’an 710032,China
  • Received:2025-11-06 Accepted:2026-01-04 Online:2026-07-28 Published:2026-07-27
  • Contact: Mu HE E-mail:394092811@qq.com

摘要:

目的 建立基于临床实验室常规检测数据的食管癌(EC)机器学习(ML)预测模型,阐明关键检测指标在模型预测中的相对贡献。 方法 采用回顾性病例对照研究,收集2019年1月- 2024年12月空军军医大学第一附属医院就诊的2 763例EC患者和3 297名年龄及性别匹配的健康对照(HC)者的临床实验室数据。对原始数据进行缺失值填充、异常值剔除及标准化预处理后,采用递归特征消除(RFE)与最小绝对收缩和选择算子(LASSO)回归联合筛选预测特征,并基于筛选特征分别构建极限梯度提升(XGBoost)、逻辑回归(LR)、轻量级梯度提升机(LightGBM)、随机森林(RF)、自适应提升算法(AdaBoost)和决策树(DT)模型。通过受试者工作特征(ROC)曲线下面积(AUC)、决策曲线分析(DCA)、校准曲线及精确率-召回率(PR)曲线对模型性能进行综合评估。采用沙普利加性解释(SHAP)方法对最优模型进行可解释性分析。 结果 从67项实验室检测指标中筛选出10项EC风险相关的预测特征,包括癌胚抗原(CEA)、尿管型定量(UCQ)、红细胞分布宽度标准差(RDW-SD)、血小板(PLT)计数、血小板分布宽度(PDW)、碱性磷酸酶(ALP)、尿酸(UA)、肌酐(Cr)、丙氨酸氨基转移酶(ALT)和总蛋白(TP)。在多模型比较中,LightGBM模型在测试集DCA中表现最优,其AUC为0.953[95%置信区间(95%CI):0.942~0.964],并在校准曲线中显示出良好的校准度和较高的临床净获益。 结论 基于临床实验室常规检测数据构建了EC的ML预测模型,并在内部数据集中完成了系统验证;其中LightGBM模型在多模型比较中表现最优,显示出较高的区分能力和良好的验证性能。

关键词: 食管肿瘤, 机器学习, 风险预测模型, 沙普利加性解释, 肿瘤筛查

Abstract:

Objective To construct a machine learning(ML)-based prediction model for esophageal cancer (EC) using routine clinical laboratory data, and to clarify the relative contribution of key laboratory indicators to model prediction. Methods A retrospective case-control study was conducted. The clinical laboratory data were collected from 2 763 patients with EC and 3 297 age-and sex-matched healthy controls(HC) treated at the First Affiliated Hospital of Air Force Medical University between January 2019 and December 2024. After missing value imputation, outlier removal, and data normalization, recursive feature elimination (RFE) combined with least absolute shrinkage and selection operator (LASSO) regression was used to screen predictive features. Based on the selected features, extreme gradient boosting (XGBoost), logistic regression (LR), light gradient boosting machine (LightGBM), random forest (RF), adaptive boosting (AdaBoost), and decision tree (DT) models were constructed. Model performance was comprehensively evaluated using the area under the receiver operating characteristic(ROC) curve (AUC), decision curve analysis (DCA), calibration curves, and precision-recall (PR) curves. SHapley Additive exPlanations (SHAP) were applied to interpret the optimal model. Results From 67 routine laboratory indicators, 10 EC-related predictive features were identified, including carcinoembryonic antigen (CEA), urinary cast quantity (UCQ), red cell distribution width-standard deviation (RDW-SD), platelet (PLT) count, platelet distribution width (PDW), alkaline phosphatase (ALP), uric acid (UA), creatinine (Cr), alanine aminotransferase (ALT), and total protein (TP). Among the compared models, the LightGBM model achieved the best performance in the test cohort, with an AUC of 0.953 [95% confidence interval(95%CI): 0.942-0.964], and demonstrated good calibration and higher clinical net benefit in calibration and DCA. Conclusion The ML-based prediction models for EC was developed using routine clinical laboratory data and systematically validated in an internal dataset. Among the compared models, the LightGBM model showed the best performance, with high discriminative ability and good validation performance.

Key words: Esophageal neoplasm, Machine learning, Risk prediction model, Shapley additive explanation, Tumor screening

中图分类号: 

  • R446