Presentation Information

[4GteX-14]Leveraging machine learning to identify amino acid sequences enhancing translation and establishing a cycle for accuracy improvement incorporating experimental results

○Gentaro Yokoyama1,2, Chie Motono1,4, Teruyo Ojima-Kato3, Hideo Nakano3, Michiaki Hamada1,2 (1. Cellular and Molecular Biotechnology Research Institute, National Institute of Advanced Industrial Science and Technology (AIST) (Japan), 2. Graduate School of Advanced Science and Engineering, Waseda University (Japan), 3. Laboratory of Molecular Biotechnology, Graduate School of Bioagricultural Sciences, Nagoya University (Japan), 4. Integrated Research Center for Self-Care Technology (IRC-SCT), National Institute of Advanced Industrial Science and Technology (AIST) (Japan))
PDF DownloadDownload PDF

Keywords:

Translation-enhancing peptide,Machine learning

[Purpose]
Advances in recombinant DNA technology have made it possible to produce arbitrary proteins in various host organisms. However, production efficiency varies considerably depending on the gene. It has been reported that introducing short peptides of a few amino acids can significantly increase protein yield. To identify novel translation-enhancing sequences and better understand the properties of known ones, this study aimed to develop an efficient computational framework using machine learning algorithms for systematic sequence exploration and iterative model refinement.
[Method]
A dataset of 158 candidate 4-amino-acid peptides was generated from E. coli in vivo experiments. Two machine learning algorithms — Random Forest (RF) and XGBoost — were trained on in vitro experimental data. Sequences were numerically represented using 157 physicochemical and geometric features derived from Z-scale, T-scale, ST-scale, VHSE-scale, and RNA structure EnsembleEnergy descriptors. Predictions were made for all ~160,000 possible 4-amino-acid combinations. A subset of top-ranked candidate sequences was selected for experimental validation, and the resulting measurements were incorporated into the training dataset, forming an active-learning cycle that was repeated for three rounds (Round 1: n=158; Round 2: n=208; Round 3: n=248).
[Results]
Analysis of in vivo data revealed that the presence of aspartic acid at the 4th position was a characteristic feature of high-translation-efficiency sequences. Feature importance analysis consistently identified Z-scale-5, T-scale-3, and EnsembleEnergy as influential variables across all rounds. The prediction accuracy of the RF model, measured by the correlation coefficient between predicted and observed values, improved progressively: 0.50, 0.64, 0.66 over the three rounds. Notably, in Round 2, the model achieved a high correlation coefficient of 0.83 with experimental values even for sequences not ranked at the top of predictions.
[Consideration]
These results suggest that translation efficiency depends not only on amino acid composition but also on position-specific effects within the peptide sequence. The consistent importance of physicochemical features such as Z-scale and T-scale across all rounds implies that the model has captured generalizable biochemical properties underlying translation enhancement. The iterative incorporation of experimental data effectively reduced prediction error and expanded the model's applicable range beyond the initial training set.
[Conclusion]
This study successfully established a machine learning-based pipeline for efficient exploration of translation-enhancing peptide sequences, coupled with an iterative loop that integrates experimental measurements to progressively improve model accuracy. Future work will focus on identifying optimal peptide insertion sites within target proteins and validating the approach in expression systems beyond E. coli.

Comment

To browse or post comments, you must log in.Log in