Back to Search
Start Over
Bayesian optimization with evolutionary and structure-based regularization for directed protein evolution
- Source :
- Algorithms for Molecular Biology, Vol 16, Iss 1, Pp 1-15 (2021), Algorithms for Molecular Biology : AMB
- Publication Year :
- 2021
- Publisher :
- BMC, 2021.
-
Abstract
- Background Directed evolution (DE) is a technique for protein engineering that involves iterative rounds of mutagenesis and screening to search for sequences that optimize a given property, such as binding affinity to a specified target. Unfortunately, the underlying optimization problem is under-determined, and so mutations introduced to improve the specified property may come at the expense of unmeasured, but nevertheless important properties (ex. solubility, thermostability, etc). We address this issue by formulating DE as a regularized Bayesian optimization problem where the regularization term reflects evolutionary or structure-based constraints. Results We applied our approach to DE to three representative proteins, GB1, BRCA1, and SARS-CoV-2 Spike, and evaluated both evolutionary and structure-based regularization terms. The results of these experiments demonstrate that: (i) structure-based regularization usually leads to better designs (and never hurts), compared to the unregularized setting; (ii) evolutionary-based regularization tends to be least effective; and (iii) regularization leads to better designs because it effectively focuses the search in certain areas of sequence space, making better use of the experimental budget. Additionally, like previous work in Machine learning assisted DE, we find that our approach significantly reduces the experimental burden of DE, relative to model-free methods. Conclusion Introducing regularization into a Bayesian ML-assisted DE framework alters the exploratory patterns of the underlying optimization routine, and can shift variant selections towards those with a range of targeted and desirable properties. In particular, we find that structure-based regularization often improves variant selection compared to unregularized approaches, and never hurts.
- Subjects :
- Mathematical optimization
Active learning
Optimization problem
Computer science
Active learning (machine learning)
QH301-705.5
Bayesian probability
QH426-470
Regularization (mathematics)
Sequence space
03 medical and health sciences
0302 clinical medicine
Structural Biology
Regularization
Genetics
Biology (General)
Molecular Biology
Selection (genetic algorithm)
030304 developmental biology
Bayesian optimization
0303 health sciences
Research
Applied Mathematics
Range (mathematics)
Computational Theory and Mathematics
Rational design
Directed evolution
Protein language model
Protein design
030217 neurology & neurosurgery
Gaussian process regression
Subjects
Details
- Language :
- English
- ISSN :
- 17487188
- Volume :
- 16
- Issue :
- 1
- Database :
- OpenAIRE
- Journal :
- Algorithms for Molecular Biology
- Accession number :
- edsair.doi.dedup.....b294c2f067db1c558bc5d1dbfdee8005