Microarray gene expression data usually consist of a large amount of genes. Among these genes, only a small fraction are informative for performing a cancer diagnostic test. This paper focuses on effective identification of informative genes. We analyze gene selection models from the perspective of optimization theory. As a result, a new strategy is designed to modify conventional search engines. Also, as overfitting is likely to occur in microarray data because of their small sample set, a point injection technique is developed to address the problem of overfitting. The proposed strategies have been evaluated on three kinds of cancer diagnosis. Our results show that the proposed strategies can improve the performance of gene selection substantially. The experimental results also indicate that the proposed methods are very robust under all of the investigated cases.
Published in:
Computational Biology and Bioinformatics, IEEE/ACM Transactions on
(Volume:4
,
Issue:
3
)
Date of Publication: July-Sept. 2007