吉林大学学报(信息科学版) ›› 2026, Vol. 44 ›› Issue (4): 889-895.

• • 上一篇    下一篇

半监督学习下医疗术语信息关键词模糊检索算法

朱越石1 , 郭凌辉2   

  1. 1. 江苏省人民医院 信息处, 南京 210029; 2. 河南大学 人工智能学院, 郑州 450046
  • 收稿日期:2025-12-27 出版日期:2026-08-06 发布日期:2026-08-07
  • 作者简介:朱越石(1990— ), 男, 南京人, 江苏省人民医院(中级)工程师, 硕士, 主要从事医学信息学、 电子信息工程研究, (Tel) 86-13951085612(E-mail)zhuyueshinj@ 163. com
  • 基金资助:
    江苏省社科应用研究精品工程重点基金资助项目(25SYA-025)

Semi-Supervised Learning-Based Fuzzy Retrieval Algorithm for Medical Term Information Keywords

ZHU Yueshi1, GUO Linghui2   

  1. 1. Information Department, Jiangsu Province Hospital, Nanjing 210029, China;2. School of Artificial Intelligence, Henan University, Zhengzhou 450046, China
  • Received:2025-12-27 Online:2026-08-06 Published:2026-08-07

摘要:

医疗领域术语具有强异构性, 且标注数据稀缺, 导致语义特征表达较弱, 进而使检索精度较低, 为此,提出了半监督学习下医疗术语信息关键词模糊检索算法。利用词向量对原始医疗术语信息进行标准化处理与映射, 基于核心元数据信息节点和用户查询信息节点, 确定检索样本之间的置信度距离, 结合向量空间模型构建医疗术语信息元数据特征空间。在元数据特征空间内, 依据元数据的时间属性, 对各类数据进行合并处理,以此同步数据索引, 并结合多级索引集群与上下文编码增强语义特征。引入基于一致性正则化与熵最小化原则的 MixMatch 半监督学习算法, 利用少量标注数据对无标签数据进行数据增强, 从而生成无标签数据的低熵伪标签, 增强语义特征表达, 并计算检索词与候选信息的相似度, 实现跨类别或近义术语的高效匹配。实验结果表明, 应用设计的方法进行医疗术语信息关键词检索, 输出的检索结果与输入关键词适配度高于 95% ,可以提高信息检索的精度。

关键词:

Abstract:

Medical terminology is characterized by strong heterogeneity and scarce labeled data, leading to weak semantic feature representation and low retrieval accuracy. To address these issues, a semi-supervised learning-based fuzzy retrieval algorithm for medical term information keywords is proposed. Word embeddings are utilized to standardize and map original medical term information. Based on core metadata information nodes and user query information nodes, the confidence distance between retrieval samples is determined, and a vector space model is employed to construct a metadata feature space for medical term information. Within this metadata feature space, data are merged according to their temporal attributes, thereby synchronizing data indexes.Semantic features are enhanced through multi-level index clustering and contextual encoding. The MixMatch semi-supervised learning algorithm, based on consistency regularization and entropy minimization principles, is introduced. A small amount of labeled data is used for data augmentation of unlabeled data, generating low-entropy pseudo-labels for unlabeled data to enhance semantic feature representation. The similarity between retrieval terms and candidate information is then calculated, enabling efficient matching across categories or synonymous terms. Experimental results demonstrate that applying the proposed method for medical term information keyword retrieval yields a retrieval result relevance exceeding 95% , effectively improving information retrieval accuracy.

Key words:

中图分类号: 

  • TP391