Back to Search Start Over

A Cross-Validated Feature Selection (CVFS) approach for extracting the most parsimonious feature sets and discovering potential antimicrobial resistance (AMR) biomarkers.

Authors :
Yang MR
Wu YW
Source :
Computational and structural biotechnology journal [Comput Struct Biotechnol J] 2022 Dec 28; Vol. 21, pp. 769-779. Date of Electronic Publication: 2022 Dec 28 (Print Publication: 2023).
Publication Year :
2022

Abstract

Understanding genes and their underlying mechanisms is critical in deciphering how antimicrobial-resistant (AMR) bacteria withstand detrimental effects of antibiotic drugs. At the same time the genes related to AMR phenotypes may also serve as biomarkers for predicting whether a microbial strain is resistant to certain antibiotic drugs. We developed a Cross-Validated Feature Selection (CVFS) approach for robustly selecting the most parsimonious gene sets for predicting AMR activities from bacterial pan-genomes. The core idea behind the CVFS approach is interrogating features among non-overlapping sub-parts of the datasets to ensure the representativeness of the features. By randomly splitting the dataset into disjoint sub-parts, conducting feature selection within each sub-part, and intersecting the features shared by all sub-parts, the CVFS approach is able to achieve the goal of extracting the most representative features for yielding satisfactory AMR activity prediction accuracy. By testing this idea on bacterial pan-genome datasets, we showed that this approach was able to extract the most succinct feature sets that predicted AMR activities very well, indicating the potential of these genes as AMR biomarkers. The functional analysis demonstrated that the CVFS approach was able to extract both known AMR genes and novel ones, suggesting the capabilities of the algorithm in selecting relevant features and highlighting the potential of the novel genes in expanding the antimicrobial resistance gene databases.<br />Competing Interests: The authors have no conflicts of interest to declare. All co-authors have seen and agree with the contents of the manuscript and there is no financial interest to report. We certify that the submission is original work and is not under review at any other publication.<br /> (© 2022 The Author(s).)

Details

Language :
English
ISSN :
2001-0370
Volume :
21
Database :
MEDLINE
Journal :
Computational and structural biotechnology journal
Publication Type :
Academic Journal
Accession number :
36698972
Full Text :
https://doi.org/10.1016/j.csbj.2022.12.046