Skip to Main Content
DNA microarray technologies are leading to an explosion in available gene expression data which simultaneously monitor the expression pattern of thousands of genes. All the genes may not be biologically significant in diagnosing the disease. In this paper, a novel approach has been proposed to select significant genes of leukemia cancer using K-Means clustering algorithm. It is an unsupervised machine learning approach, which is being used to identify the unknown patterns from the huge amount of data. The proposed K-Means algorithm has been experimented to cluster the genes for K=5,10 and 15. The significant genes have been identified through the best accuracy obtained from the clusters generated. The accuracy of the clusters are determined again by using K-Means algorithm compared with ground truth values.