吉林大学学报(信息科学版) ›› 2026, Vol. 44 ›› Issue (4): 956-962.

• • 上一篇    下一篇

混合语义相似度提取下高维重叠数据智能模糊检索算法

沈飞洋, 刘 涛   

  1. 杭州电子科技大学 法学院, 杭州 310018
  • 收稿日期:2025-06-25 出版日期:2026-08-06 发布日期:2026-08-06
  • 作者简介:沈飞洋(1999— ), 男, 杭州人, 杭州电子科技大学硕士研究生, 主要从事人工智能研究, (Tel)86-19858190382(E-mail)Shenfeiyang0606@ 163. com; 刘涛(1982— ), 男, 河南尉氏人, 杭州电子科技大学副教授, 硕士生导师, 主要从事社会人类学与人工智能研究, (Tel)86-19858190382(E-mail)Shenfeiyang0606@ 163. com。
  • 基金资助:
    浙江省自然科学基金资助项目(POF1D6215623)

Intelligent Fuzzy Retrieval Algorithm for High Dimensional Overlapping Data under Hybrid Semantic Similarity Extraction

SHEN Feiyang, LIU Tao   

  1. School of Law, Hangzhou Dianzi University, Hangzhou 310018, China
  • Received:2025-06-25 Online:2026-08-06 Published:2026-08-06

摘要:

针对在数据智能检索过程中, 单一语义相似度的局限性可能会影响数据相关性的判断, 导致检索结果的排序敏感度不高的问题, 提出了混合语义相似度提取下高维重叠数据智能模糊检索算法。建立高维重叠数据的信息粒度空间, 通过密度聚类与模糊度量, 识别出空间中的重叠数据所在的区域, 从而在不改变数据结构的前提下, 实现高维数据降维。混合均值函数、近似线性化统计、规则方法 3 种语义相似计算方法, 综合分析降维数据的相关性。在最近邻算法的遍历下, 按序提取出语义相似度较高的数据信息, 生成数据检索列表。算例测试结果表明, 该算法所表现出的 NDCG(Normalized Discounted Cumulative Gain)指标值平均为 0. 867, 检索结果具有较高的排序敏感度, 检索质量更高, 具有良好的实践应用前景。

关键词:

Abstract:

In the process of intelligent data retrieval, the limitation of a single semantic similarity may affect the judgment of data relevance, resulting in a low ranking sensitivity of the retrieval results. To alleviate this problem, an intelligent fuzzy retrieval algorithm is proposed for high-dimensional overlapping data under hybrid semantic similarity extraction. The information granularity space of high-dimensional overlapping data is established. Through density clustering and fuzzy measurement, the regions where the overlapping data is located in the space are identified, thereby achieving the dimension reduction of high-dimensional data without changing the data structure. Three semantic similarity calculation methods, namely the mixed mean function, approximate inearization statistics, and rule method, are used to comprehensively analyze the correlation of dimensionality reduction data. Under the traversal of the nearest neighbor algorithm, the data information with higher semantic similarity is extracted in sequence to generate the data retrieval list. The results of the case study show that the average NDCG(Normalized Discounted Cumulative Gain) index value exhibited by this algorithm is 0. 867. The retrieval results have a high ranking sensitivity and higher retrieval quality, and have a good practical application prospect.

Key words:

中图分类号: 

  • TP391