吉林大学学报(工学版) ›› 2013, Vol. 43 ›› Issue (01): 130-134.

• 论文 • 上一篇    下一篇

混合属性数据聚类的新方法

白天1,2, 冀进朝1, 何加亮1, 周春光1   

  1. 1. 吉林大学 计算机科学与技术学院, 长春 130012;
    2. 新泽西州立大学 计算机系, 新泽西州, NJ 08901
  • 收稿日期:2011-12-28 出版日期:2013-01-01 发布日期:2013-01-01
  • 通讯作者: 周春光(1947-),男,教授,博士生导师.研究方向:计算智能.E-mail:cgzhou@jlu.edu.cn E-mail:cgzhou@jlu.edu.cn
  • 作者简介:白天(1983-),男,讲师,博士.研究方向:数据挖掘,生物信息学.E-mail:baitian09@mails.jlu.edu.cn
  • 基金资助:

    国家自然科学基金项目(61175023,60973092,60903097);符号计算与知识工程教育部重点实验室项目;国家留学基金委项目(2010617098).

New clustering method of mixed-attribute data

BAI Tian1,2, JI Jin-chao1, HE Jia-liang1, ZHOU Chun-guang1   

  1. 1. College of Computer Science and Technology, Jilin University, Changchun 130022, China;
    2. Department of Computer Science, Rutgers, the State University of New Jersey, New Jersey, NJ 08901, USA
  • Received:2011-12-28 Online:2013-01-01 Published:2013-01-01

摘要: 提出了一种数值型和类别型混合属性数据聚类的全局算法。算法通过随机选取足够多的初始原型来覆盖数据集的全局分布信息,然后通过评估函数迭代地消去多余的原型。最后对本文算法进行了验证,证明了该算法的有效性和收敛性。并与其他已有同类型算法的聚类结果进行比较,说明本文算法对混合属性数据具有更高的聚类准确度,为解决混合型数据聚类问题提供了一种新途径。

关键词: 人工智能, 数据聚类, 数据挖掘, K原型算法, 混合属性数据

Abstract: A new Global k-Prototype (GKP) algorithm is proposed for clustering mixed numeric and categorical data. First, the algorithm randomly selects a sufficiently large number of initial prototypes to account for the global distribution of the data sets. Then, it progressively eliminates the redundant prototypes using an iterative optimization process with an elimination criterion function. Systematic experiments were carried out with data from widely used datasets in this area. Experimental results and comparative evaluation show the high performance and consistency of the proposed algorithm. Compared with other well-known mixed data clustering algorithms, the proposed algorithm significantly improves the clustering accuracy.

Key words: artificial intelligence, clustering, data mining, K-prototypes algorithm, mixed attribute data

中图分类号: 

  • TP181
[1] Jain A K, Murty M N, Flynn P J. Data clustering: a review[J]. ACM Computing Surveys, 1999, 31(3): 264-323.

[2] 徐森, 卢志茂, 顾国昌. 结合K均值和非负矩阵分解集成文本聚类算法[J]. 吉林大学学报:工学版, 2011,41(4): 1077-1082. Xu Sen, Lu Zhi-mao, Gu Guo-chang. Integrating K-means and non-negative matrix factorization to ensemble document clustering[J]. Journal of Jilin University(Engineering and Technology Edition), 2011,41(4): 1077-1082.

[3] Han J, Kamber M. Data Mining Concepts and Techniques[M]. San Francisco: Morgan Kaufmann, 2001.

[4] MacQueen J. Some methods for classification and analysis of multivariate observation//Proc 5th Berkeley Symp on Mathematical Statistics and Probability, 1967: 281-297.

[5] Anderberg Michael R. Cluster Analysis for Applications[M]. New York: Academic Press, 1973.

[6] Hsu C C, Huang Y P. Incremental clustering of mixed data based on distance hierarchy[J]. Expert Systems with Applications, 2008, 35(3): 1177-1185.

[7] Huang Z. Clustering large data sets with mixed numeric and categorical values//Proceedings of the First Pacific Asia Knowledge Discovery and Data Mining Conference, World Scientific, Singapore, 1997:21-34.

[8] Bezdek J C, Keller J, Krisnapuram R. Fuzzy Models and Algorithms for Pattern Recognition and Image Processing[M]. Boston: Kluwer Academy Publishers, 1999.

[9] Ahmad A, Dey L. Algorithm for fuzzy clustering of mixed data with numeric and categorical attributes[J]. LNCS, 2005, 3816: 561-572.

[10] Chatzis Sotirios P. A fuzzy c-means-type algorithm for clustering of data with mixed numeric and categorical attributes employing a probabilistic dissimilarity functional[J]. Expert Systems with Applications, 2011, 38(7): 8684-8689.

[11] Zheng Z, Gong M G, Ma J J,et al. Unsupervised evolutionary clustering algorithm for mixed type data//IEEE Congress on Evolutionary Computation, 2010.

[12] Li C, Biswas G. Unsupervised learning with mixed numeric and nominal data//IEEE Transactions on Knowledge and Data Engineering, 2002, 14(4): 673-690.

[13] Hsu C C, Chen Y C. Mining of mixed data with application to catalog marketing[J]. Expert Systems with Applications, 2007, 32(1): 12-27.

[14] Ahmad A, Dey L. A k-mean clustering algorithm for mixed numeric and categorical data[J]. Data & Knowledge Engineering, 2007, 63(2): 503-527.

[15] Merz C, Murphy P, Aha D. UCI repository of machine learning databases. Irvine: Department of Information and Computer Science, University of California, 1997.

[16] Huang Z, Ng M K. A fuzzy k-modes algorithm for clustering categorical data[J]. IEEE Trans Fuzzy System, 1999, 7(4): 446-452.
[1] 董飒, 刘大有, 欧阳若川, 朱允刚, 李丽娜. 引入二阶马尔可夫假设的逻辑回归异质性网络分类方法[J]. 吉林大学学报(工学版), 2018, 48(5): 1571-1577.
[2] 顾海军, 田雅倩, 崔莹. 基于行为语言的智能交互代理[J]. 吉林大学学报(工学版), 2018, 48(5): 1578-1585.
[3] 王旭, 欧阳继红, 陈桂芬. 基于垂直维序列动态时间规整方法的图相似度度量[J]. 吉林大学学报(工学版), 2018, 48(4): 1199-1205.
[4] 张浩, 占萌苹, 郭刘香, 李誌, 刘元宁, 张春鹤, 常浩武, 王志强. 基于高通量数据的人体外源性植物miRNA跨界调控建模[J]. 吉林大学学报(工学版), 2018, 48(4): 1206-1213.
[5] 黄岚, 纪林影, 姚刚, 翟睿峰, 白天. 面向误诊提示的疾病-症状语义网构建[J]. 吉林大学学报(工学版), 2018, 48(3): 859-865.
[6] 李雄飞, 冯婷婷, 骆实, 张小利. 基于递归神经网络的自动作曲算法[J]. 吉林大学学报(工学版), 2018, 48(3): 866-873.
[7] 刘杰, 张平, 高万夫. 基于条件相关的特征选择方法[J]. 吉林大学学报(工学版), 2018, 48(3): 874-881.
[8] 邓剑勋, 熊忠阳, 邓欣. 基于谱聚类矩阵的改进DNALA算法[J]. 吉林大学学报(工学版), 2018, 48(3): 903-908.
[9] 王旭, 欧阳继红, 陈桂芬. 基于多重序列所有公共子序列的启发式算法度量多图的相似度[J]. 吉林大学学报(工学版), 2018, 48(2): 526-532.
[10] 杨欣, 夏斯军, 刘冬雪, 费树岷, 胡银记. 跟踪-学习-检测框架下改进加速梯度的目标跟踪[J]. 吉林大学学报(工学版), 2018, 48(2): 533-538.
[11] 刘雪娟, 袁家斌, 许娟, 段博佳. 量子k-means算法[J]. 吉林大学学报(工学版), 2018, 48(2): 539-544.
[12] 曲慧雁, 赵伟, 秦爱红. 基于优化算子的快速碰撞检测算法[J]. 吉林大学学报(工学版), 2017, 47(5): 1598-1603.
[13] 李嘉菲, 孙小玉. 基于谱分解的不确定数据聚类方法[J]. 吉林大学学报(工学版), 2017, 47(5): 1604-1611.
[14] 邵克勇, 陈丰, 王婷婷, 王季驰, 周立朋. 无平衡点分数阶混沌系统全状态自适应控制[J]. 吉林大学学报(工学版), 2017, 47(4): 1225-1230.
[15] 王生生, 王创峰, 谷方明. OPRA方向关系网络的时空推理[J]. 吉林大学学报(工学版), 2017, 47(4): 1238-1243.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!