Presentation Information
[P01-032]Leveraging high-throughput sequencing and deep learning to predict high lipid-producing variants in Yarrowia lipolytica
○Jaehoon Kim1, Min Kyeong Kim2, Sungmin Hwang3, Yongjae Lee4, Jiwon Kim1, Ji Hyun Jung1, Kyung-Jin Kim5, Hyeoncheol Francis Son6, Sun-Mi Lee4, Daechan Park1,7 (1. Department of Molecular Science and Technology, Ajou Univ., Suwon 16499, Republic of Korea (Korea), 2. Clean Energy Research Center, Korea Institute of Science and Technology (KIST), Seoul, 02792, Republic of Korea (Korea), 3. Division of Convergence on Marine Science, Korea Maritime & Ocean Univ., Busan 49112, Republic of Korea (Korea), 4. Department of Environmental Science and Ecological Engineering, Korea Univ., Seoul 02841, Republic of Korea (Korea), 5. KNU Institute for Microorganisms, Kyungpook National Univ., Daegu, Republic of Korea (Korea), 6. School of Biological Sciences and Technology, Chonnam National Univ., Gwangju 61186, Republic of Korea (Korea), 7. Advanced College of Bio-Convergence Engineering, Ajou Univ., Suwon 16499, Republic of Korea (Korea))
Keywords:
Enzyme engineering,Directed evolution,High-throughput sequencing,Deep learning,Synthetic biology
Exploring the vast sequence space in directed evolution remains challenging, as library diversity increases. Although high-throughput sequencing (HTS) provides large amounts of sequence data, it is insufficient to fully cover diversity, which grows exponentially with the number of targeted positions. To address this limitation, computational approaches such as deep learning (DL) are needed to model sequence-fitness relationships and identify high-performing variants beyond the observed space.
To improve lipid production in Yarrowia lipolytica, we targeted the enzymes MCE2 and ZWF1 involved in NADPH regeneration. Site-saturation mutagenesis libraries were constructed for four regions in each enzyme and screened over four rounds by fluorescence-activated cell sorting (FACS). After screening, HTS was performed to quantify clones and their enrichment. Unique molecular identifiers (UMIs) were introduced to reduce amplification bias and sequencing artifacts. Clonal diversity showed that directed evolution converged by the second round, after which the diversity stabilized. By the final round, some regions were highly polarized with a single dominant clone exceeding 90%, whereas others retained diverse populations. These results indicate that each target region followed a distinct evolutionary trajectory.
Despite broad HTS coverage, most clones remained unobserved. To estimate enzyme fitness for unseen clones, multiple deep learning (DL) models, including convolutional neural networks (CNNs), long short-term memory networks (LSTMs), and transformers, were trained on HTS-derived sequence–enrichment data. These models achieved moderate predictive performance on the test set (Pearson’s r2 = 0.53 ~ 0.57). Despite using the same training set, they exhibited distinct sequence preferences, implying that each captures different features of the fitness landscape. To mitigate model-specific bias, a consensus score was computed by averaging normalized outputs across models.
We further explored unobserved sequence space by generating in silico libraries using three strategies. First, top-ranked clones were used as a backbone, and combinatorial libraries were generated by combining high-frequency residues at each position. Second, probabilistic sampling was applied to incorporate lower-frequency residues around this backbone. Third, stepwise mutational walks were simulated from the template sequence by accumulating mutations. Despite these distinct approaches, top-ranked clones converged on similar amino acid motifs. Notably, models prioritized a stop codon at a specific region despite its absence in training set, suggesting extrapolation capability beyond experimentally observed variants. Overall, integrating HTS with DL expands searchable sequence space and facilitates identification of high lipid-producing variants beyond conventional screening.
To improve lipid production in Yarrowia lipolytica, we targeted the enzymes MCE2 and ZWF1 involved in NADPH regeneration. Site-saturation mutagenesis libraries were constructed for four regions in each enzyme and screened over four rounds by fluorescence-activated cell sorting (FACS). After screening, HTS was performed to quantify clones and their enrichment. Unique molecular identifiers (UMIs) were introduced to reduce amplification bias and sequencing artifacts. Clonal diversity showed that directed evolution converged by the second round, after which the diversity stabilized. By the final round, some regions were highly polarized with a single dominant clone exceeding 90%, whereas others retained diverse populations. These results indicate that each target region followed a distinct evolutionary trajectory.
Despite broad HTS coverage, most clones remained unobserved. To estimate enzyme fitness for unseen clones, multiple deep learning (DL) models, including convolutional neural networks (CNNs), long short-term memory networks (LSTMs), and transformers, were trained on HTS-derived sequence–enrichment data. These models achieved moderate predictive performance on the test set (Pearson’s r2 = 0.53 ~ 0.57). Despite using the same training set, they exhibited distinct sequence preferences, implying that each captures different features of the fitness landscape. To mitigate model-specific bias, a consensus score was computed by averaging normalized outputs across models.
We further explored unobserved sequence space by generating in silico libraries using three strategies. First, top-ranked clones were used as a backbone, and combinatorial libraries were generated by combining high-frequency residues at each position. Second, probabilistic sampling was applied to incorporate lower-frequency residues around this backbone. Third, stepwise mutational walks were simulated from the template sequence by accumulating mutations. Despite these distinct approaches, top-ranked clones converged on similar amino acid motifs. Notably, models prioritized a stop codon at a specific region despite its absence in training set, suggesting extrapolation capability beyond experimentally observed variants. Overall, integrating HTS with DL expands searchable sequence space and facilitates identification of high lipid-producing variants beyond conventional screening.
Comment
To browse or post comments, you must log in.Log in
