IDEAS home Printed from https://ideas.repec.org/a/plo/pone00/0214406.html
   My bibliography  Save this article

Classification of high dimensional biomedical data based on feature selection using redundant removal

Author

Listed:
  • Bingtao Zhang
  • Peng Cao

Abstract

High dimensional biomedical data contain tens of thousands of features, accurate and effective identification of the core features in these data can be used to assist diagnose related diseases. However, there are often a large number of irrelevant or redundant features in biomedical data, which seriously affect subsequent classification accuracy and machine learning efficiency. To solve this problem, a novel filter feature selection algorithm based on redundant removal (FSBRR) is proposed to classify high dimensional biomedical data in this paper. First of all, two redundant criteria are determined by vertical relevance (the relationship between feature and class attribute) and horizontal relevance (the relationship between feature and feature). Secondly, to quantify redundant criteria, an approximate redundancy feature framework based on mutual information (MI) is defined to remove redundant and irrelevant features. To evaluate the effectiveness of our proposed algorithm, controlled trials based on typical feature selection algorithm are conducted using three different classifiers, and the experimental results indicate that the FSBRR algorithm can effectively reduce the feature dimension and improve the classification accuracy. In addition, an experiment of small sample dataset is designed and conducted in the section of discussion and analysis to clarify the specific implementation process of FSBRR algorithm more clearly.

Suggested Citation

  • Bingtao Zhang & Peng Cao, 2019. "Classification of high dimensional biomedical data based on feature selection using redundant removal," PLOS ONE, Public Library of Science, vol. 14(4), pages 1-19, April.
  • Handle: RePEc:plo:pone00:0214406
    DOI: 10.1371/journal.pone.0214406
    as

    Download full text from publisher

    File URL: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0214406
    Download Restriction: no

    File URL: https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0214406&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pone.0214406?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    References listed on IDEAS

    as
    1. Chen Zhang & Zhiwei Ni & Liping Ni & Na Tang, 2016. "Feature selection method based on multi-fractal dimension and harmony search algorithm and its application," International Journal of Systems Science, Taylor & Francis Journals, vol. 47(14), pages 3476-3486, October.
    2. Hossam M Zawbaa & E Emary & Crina Grosan, 2016. "Feature Selection via Chaotic Antlion Optimization," PLOS ONE, Public Library of Science, vol. 11(3), pages 1-21, March.
    3. Olvi L. Mangasarian & W. Nick Street & William H. Wolberg, 1995. "Breast Cancer Diagnosis and Prognosis Via Linear Programming," Operations Research, INFORMS, vol. 43(4), pages 570-577, August.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Akampurira Paul & Mutebi Joe & Mugisha Brian & Muhaise Hussein & Kyomuhangi Rosette, 2024. "Exploring Dimensionality Reduction Techniques for Improved Breast Cancer Diagnosis," International Journal of Research and Scientific Innovation, International Journal of Research and Scientific Innovation (IJRSI), vol. 11(5), pages 808-824, May.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Sexton, Randall S. & Dorsey, Robert E. & Johnson, John D., 1999. "Optimization of neural networks: A comparative analysis of the genetic algorithm and simulated annealing," European Journal of Operational Research, Elsevier, vol. 114(3), pages 589-601, May.
    2. Brandner, Hubertus & Lessmann, Stefan & Voß, Stefan, 2013. "A memetic approach to construct transductive discrete support vector machines," European Journal of Operational Research, Elsevier, vol. 230(3), pages 581-595.
    3. W. Art Chaovalitwongse & Ya-Ju Fan & Rajesh C. Sachdeo, 2008. "Novel Optimization Models for Abnormal Brain Activity Classification," Operations Research, INFORMS, vol. 56(6), pages 1450-1460, December.
    4. Tamilselvan, Prasanna & Wang, Pingfeng, 2013. "Failure diagnosis using deep belief learning based health state classification," Reliability Engineering and System Safety, Elsevier, vol. 115(C), pages 124-135.
    5. Yaqiong Cui & Jukka Sirén & Timo Koski & Jukka Corander, 2016. "Simultaneous Predictive Gaussian Classifiers," Journal of Classification, Springer;The Classification Society, vol. 33(1), pages 73-102, April.
    6. Ramazan Ünlü & Petros Xanthopoulos, 2019. "A weighted framework for unsupervised ensemble learning based on internal quality measures," Annals of Operations Research, Springer, vol. 276(1), pages 229-247, May.
    7. Morris, Katherine & McNicholas, Paul D., 2016. "Clustering, classification, discriminant analysis, and dimension reduction via generalized hyperbolic mixtures," Computational Statistics & Data Analysis, Elsevier, vol. 97(C), pages 133-150.
    8. Ryu, Young U. & Chandrasekaran, R. & Jacob, Varghese S., 2007. "Breast cancer prediction using the isotonic separation technique," European Journal of Operational Research, Elsevier, vol. 181(2), pages 842-854, September.
    9. B Baesens & C Mues & D Martens & J Vanthienen, 2009. "50 years of data mining and OR: upcoming trends and challenges," Journal of the Operational Research Society, Palgrave Macmillan;The OR Society, vol. 60(1), pages 16-23, May.
    10. Sahin, Özge & Czado, Claudia, 2022. "Vine copula mixture models and clustering for non-Gaussian data," Econometrics and Statistics, Elsevier, vol. 22(C), pages 136-158.
    11. A. Astorino & M. Gaudioso, 2002. "Polyhedral Separability Through Successive LP," Journal of Optimization Theory and Applications, Springer, vol. 112(2), pages 265-293, February.
    12. Alejandro Murua & Nicolas Wicker, 2015. "Kernel-based mixture models for classification," Computational Statistics, Springer, vol. 30(2), pages 317-344, June.
    13. Pedro Duarte Silva, A., 2017. "Optimization approaches to Supervised Classification," European Journal of Operational Research, Elsevier, vol. 261(2), pages 772-788.
    14. Sung, Bongjung & Lee, Jaeyong, 2023. "Covariance structure estimation with Laplace approximation," Journal of Multivariate Analysis, Elsevier, vol. 198(C).
    15. Wang, Wan-Lun, 2015. "Mixtures of common t-factor analyzers for modeling high-dimensional data with missing values," Computational Statistics & Data Analysis, Elsevier, vol. 83(C), pages 223-235.
    16. Wang, Haifeng & Zheng, Bichen & Yoon, Sang Won & Ko, Hoo Sang, 2018. "A support vector machine-based ensemble algorithm for breast cancer diagnosis," European Journal of Operational Research, Elsevier, vol. 267(2), pages 687-699.
    17. Jun-Ya Gotoh & Michael Jong Kim & Andrew E. B. Lim, 2017. "Calibration of Distributionally Robust Empirical Optimization Models," Papers 1711.06565, arXiv.org, revised May 2020.
    18. Xin Liu & Bangxin Zhao & Wenqing He, 2020. "Simultaneous Feature Selection and Classification for Data-Adaptive Kernel-Penalized SVM," Mathematics, MDPI, vol. 8(10), pages 1-22, October.
    19. Eva K. Lee & Richard J. Gallagher & David A. Patterson, 2003. "A Linear Programming Approach to Discriminant Analysis with a Reserved-Judgment Region," INFORMS Journal on Computing, INFORMS, vol. 15(1), pages 23-41, February.
    20. W. N. Street & O. L. Mangasarian, 1998. "Improved Generalization via Tolerant Training," Journal of Optimization Theory and Applications, Springer, vol. 96(2), pages 259-279, February.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pone00:0214406. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosone (email available below). General contact details of provider: https://journals.plos.org/plosone/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.