Classification across gene expression microarray studiesReport as inadecuate




Classification across gene expression microarray studies - Download this document for free, or read online. Document in PDF available to download.

BMC Bioinformatics

, 10:453

First Online: 30 December 2009Received: 08 September 2008Accepted: 30 December 2009

Abstract

BackgroundThe increasing number of gene expression microarray studies represents an important resource in biomedical research. As a result, gene expression based diagnosis has entered clinical practice for patient stratification in breast cancer. However, the integration and combined analysis of microarray studies remains still a challenge. We assessed the potential benefit of data integration on the classification accuracy and systematically evaluated the generalization performance of selected methods on four breast cancer studies comprising almost 1000 independent samples. To this end, we introduced an evaluation framework which aims to establish good statistical practice and a graphical way to monitor differences. The classification goal was to correctly predict estrogen receptor status negative-positive and histological grade low-high of each tumor sample in an independent study which was not used for the training. For the classification we chose support vector machines SVM, predictive analysis of microarrays PAM, random forest RF and k-top scoring pairs kTSP. Guided by considerations relevant for classification across studies we developed a generalization of kTSP which we evaluated in addition. Our derived version DV aims to improve the robustness of the intrinsic invariance of kTSP with respect to technologies and preprocessing.

ResultsFor each individual study the generalization error was benchmarked via complete cross-validation and was found to be similar for all classification methods. The misclassification rates were substantially higher in classification across studies, when each single study was used as an independent test set while all remaining studies were combined for the training of the classifier. However, with increasing number of independent microarray studies used in the training, the overall classification performance improved. DV performed better than the average and showed slightly less variance. In particular, the better predictive results of DV in across platform classification indicate higher robustness of the classifier when trained on single channel data and applied to gene expression ratios.

ConclusionsWe present a systematic evaluation of strategies for the integration of independent microarray studies in a classification task. Our findings in across studies classification may guide further research aiming on the construction of more robust and reliable methods for stratification and diagnosis in clinical practice.

Electronic supplementary materialThe online version of this article doi:10.1186-1471-2105-10-453 contains supplementary material, which is available to authorized users.

Download fulltext PDF



Author: Andreas Buness - Markus Ruschhaupt - Ruprecht Kuner - Achim Tresch

Source: https://link.springer.com/







Related documents