Presentation Information
[2Biocat-03]Navigating protein fitness landscapes with small-data supervised learning: Accelerating multi-objective protein engineering
○Mitsuo Umetsu1,2,3 (1. Tohoku University (Japan), 2. RIKEN (Japan), 3. RevolKa Ltd. (Japan))
Keywords:
Protein engineering,Machine learning,molecular evolution
The rapid development of high-performance enzymes is a central challenge in modern biotechnology and sustainable bioprocessing. However, conventional directed evolution often suffers from a severe “screening bottleneck,” requiring the evaluation of massive mutant libraries to identify improved variants.
In this presentation, I will present a machine-learning–guided strategy that enables efficient navigation of protein fitness landscapes using surprisingly small experimental datasets. By integrating supervised learning with carefully designed mutational libraries, we demonstrate that predictive models can be trained with as few as ~100 experimentally characterized variants. These models allow the identification of highly improved protein variants while dramatically reducing the scale of experimental screening.
While machine learning in biotechnology is often associated with large datasets, our results show that high-quality experimental measurements from a limited number of variants can provide sufficient information to guide protein evolution. This data-efficient strategy enables a shift from brute-force screening to data-driven design, making advanced protein engineering accessible even without ultra-high-throughput screening platforms.
Importantly, this framework is particularly powerful for multi-objective optimization, where proteins must simultaneously satisfy multiple functional requirements such as catalytic activity, stability, and substrate specificity. Because each iteration requires only ~100 measurements, multiple phenotypic traits can be experimentally characterized in parallel, enabling the identification of Pareto-optimal protein variants that balance competing functional constraints.
I will discuss the design principles underlying this ML-guided evolution strategy, including training data selection, model construction, and iterative optimization cycles. Applications to protein engineering demonstrate how this approach can rapidly generate high-performance biocatalysts.
In this presentation, I will present a machine-learning–guided strategy that enables efficient navigation of protein fitness landscapes using surprisingly small experimental datasets. By integrating supervised learning with carefully designed mutational libraries, we demonstrate that predictive models can be trained with as few as ~100 experimentally characterized variants. These models allow the identification of highly improved protein variants while dramatically reducing the scale of experimental screening.
While machine learning in biotechnology is often associated with large datasets, our results show that high-quality experimental measurements from a limited number of variants can provide sufficient information to guide protein evolution. This data-efficient strategy enables a shift from brute-force screening to data-driven design, making advanced protein engineering accessible even without ultra-high-throughput screening platforms.
Importantly, this framework is particularly powerful for multi-objective optimization, where proteins must simultaneously satisfy multiple functional requirements such as catalytic activity, stability, and substrate specificity. Because each iteration requires only ~100 measurements, multiple phenotypic traits can be experimentally characterized in parallel, enabling the identification of Pareto-optimal protein variants that balance competing functional constraints.
I will discuss the design principles underlying this ML-guided evolution strategy, including training data selection, model construction, and iterative optimization cycles. Applications to protein engineering demonstrate how this approach can rapidly generate high-performance biocatalysts.
Comment
To browse or post comments, you must log in.Log in
