Pathway-BasedFeature Selection Algorithm for Cancer Microarray Data
[摘要] Classification of cancers based on gene expressions produces better accuracywhen compared to that of the clinical markers. Feature selection improvesthe accuracy of these classification algorithms by reducing the chanceof overfitting that happens due to large number of features. We develop anew feature selection method calledBiological Pathway-based Feature Selection (BPFS)for microarray data. Unlike most of the existing methods,our method integrates signaling and gene regulatory pathways with geneexpression data to minimize the chance of overfitting of the method and toimprove the test accuracy. Thus, BPFS selects a biologically meaningful featureset that is minimally redundant. Our experiments on published breastcancer datasets demonstrate that all of the top 20 genes found by our methodare associated with cancer. Furthermore, the classification accuracy of oursignature is up to 18% better than that of vant Veers 70 gene signature,and it is up to 8% better accuracy than the best published feature selectionmethod, I-RELIEF.
[发布日期] [发布机构]
[效力级别] [学科分类] 生物技术
[关键词] [时效性]