Presentation Information
[P01-026]Machine Learning–Guided Fragment Recombination for Simultaneous Optimization of Multiple Properties in Enzyme Engineering
○Yuto Kubo1, Hikaru Nakazawa1, Tomoyuki Ito1, Ryo Yamazaki2, Mitsuo Umetsu1,2 (1. Tohoku Univ. (Japan), 2. RevolKa Ltd. (Japan))
Keywords:
Enzyme Engineering,Machine Learning,Fragment Recombination,Chimeric mutant
[Purpose]
Recent advances in machine learning (ML) have enabled enzyme engineering from small datasets; however, exhaustive exploration of sequence space remains computationally prohibitive.
In this study, we aimed to design a mutation space that enriches promising variants while remaining computationally tractable by introducing a fragment-based strategy that incorporates multi-residue interactions (epistasis).
[Method]
We used TEM-1 β-lactamase and its ancestral sequence GNCA (Gram-Negative Common Ancestor) as parental enzymes. Sequence regions with multiple substitutions were defined as “fragments,” and 10 fragments were designed using the Schema algorithm, yielding 210 (1,024) variants.
Chimeric mutants were generated by fragment-level PCR recombination. Catalytic activity was measured using benzylpenicillin, and thermal stability was evaluated as melting temperature (Tm) by differential scanning calorimetry (DSC).
Experimental data were used to train an ML model, followed by Bayesian optimization to predict high-performing variants, of which the top 10 were synthesized and validated.
[Results]
42 chimeric variants were obtained from the PCR-constructed library. A binomial test (α = 0.05) showed no significant bias in parental sequence distribution in 8 of 10 fragments, indicating near-random sampling across fragment positions. Among the 42 variants, 38 (90%) retained catalytic activity. Catalytic activity increased up to ∼1.7-fold relative to TEM-1, whereas no variants exceeded the thermal stability of GNCA (ΔTm ≦ −4.13 ℃). Variants with higher thermal stability tended to show lower activity, indicating a trade-off, although some variants retained high activity with moderate stability. An ML model was trained on the dataset, and Bayesian optimization was applied using a composite score favoring balance between activity and stability. Among the top 10 predicted variants, 2 were included in the training dataset. Top candidates showed a shift toward variants balancing activity and stability, with improved score distribution, and the best variant achieved a score of 0.58, exceeding the highest training score (0.41) by ∼42%.
[Consideration]
The fragment-based design captured epistatic interactions while reducing structural disruption, enabling efficient exploration of sequence space.
The high retention of functional variants contributed to effective model training, highlighting the importance of dataset quality in small-data learning.
This approach also enables efficient sampling of epistatic combinations that are difficult to access by conventional residue-level mutagenesis.
[Conclusion]
We developed a fragment-based recombination strategy integrated with ML for efficient enzyme optimization from small datasets. This approach constructs compact, functionally enriched libraries and enables the identification of variants with balanced properties. It provides a general framework for sequence space design in ML-guided protein engineering.
Recent advances in machine learning (ML) have enabled enzyme engineering from small datasets; however, exhaustive exploration of sequence space remains computationally prohibitive.
In this study, we aimed to design a mutation space that enriches promising variants while remaining computationally tractable by introducing a fragment-based strategy that incorporates multi-residue interactions (epistasis).
[Method]
We used TEM-1 β-lactamase and its ancestral sequence GNCA (Gram-Negative Common Ancestor) as parental enzymes. Sequence regions with multiple substitutions were defined as “fragments,” and 10 fragments were designed using the Schema algorithm, yielding 210 (1,024) variants.
Chimeric mutants were generated by fragment-level PCR recombination. Catalytic activity was measured using benzylpenicillin, and thermal stability was evaluated as melting temperature (Tm) by differential scanning calorimetry (DSC).
Experimental data were used to train an ML model, followed by Bayesian optimization to predict high-performing variants, of which the top 10 were synthesized and validated.
[Results]
42 chimeric variants were obtained from the PCR-constructed library. A binomial test (α = 0.05) showed no significant bias in parental sequence distribution in 8 of 10 fragments, indicating near-random sampling across fragment positions. Among the 42 variants, 38 (90%) retained catalytic activity. Catalytic activity increased up to ∼1.7-fold relative to TEM-1, whereas no variants exceeded the thermal stability of GNCA (ΔTm ≦ −4.13 ℃). Variants with higher thermal stability tended to show lower activity, indicating a trade-off, although some variants retained high activity with moderate stability. An ML model was trained on the dataset, and Bayesian optimization was applied using a composite score favoring balance between activity and stability. Among the top 10 predicted variants, 2 were included in the training dataset. Top candidates showed a shift toward variants balancing activity and stability, with improved score distribution, and the best variant achieved a score of 0.58, exceeding the highest training score (0.41) by ∼42%.
[Consideration]
The fragment-based design captured epistatic interactions while reducing structural disruption, enabling efficient exploration of sequence space.
The high retention of functional variants contributed to effective model training, highlighting the importance of dataset quality in small-data learning.
This approach also enables efficient sampling of epistatic combinations that are difficult to access by conventional residue-level mutagenesis.
[Conclusion]
We developed a fragment-based recombination strategy integrated with ML for efficient enzyme optimization from small datasets. This approach constructs compact, functionally enriched libraries and enables the identification of variants with balanced properties. It provides a general framework for sequence space design in ML-guided protein engineering.
Comment
To browse or post comments, you must log in.Log in
