Presentation Information

[P04-459]Benchmarking Large Language Models for Interpreting Genome-scale Metabolic Models for Pathway/Strain Engineering

○Jing Wui Yeoh1,2, Chinari Pawan Kumar Patro1,2, Lim Soon Wong3, Chueh Loo Poh1,2,4,5 (1. National Centre for Engineering Biology (NCEB) (Singapore), 2. NUS Synthetic Biology for Clinical and Technological Innovation (SynCTI), National University of Singapore, 117456 (Singapore), 3. School of Computing, National University of Singapore, 13 Computing Drive, Singapore 117417 (Singapore), 4. Department of Biomedical Engineering, College of Design and Engineering, National University of Singapore (Singapore), 5. Department of Biochemistry, School of Medicine, National University of Singapore (Singapore))
PDF DownloadDownload PDF

Keywords:

Genome-scale metabolic model,LLM benchmarking,metabolic flux prediction,strain engineering,pathway engineering

Genome-scale metabolic models (GSM) are essential tools for pathway and strain engineering, enabling systematic interventions in cellular metabolism to optimise bioproduction. GSMs support constraint-based analyses, such as flux balance analysis, which inform pathway reconstruction, knockout analysis and other engineering strategies to guide experimental design to improve yields of target chemicals. However, implementing efficient workflows for GSM analysis remains technically demanding, requiring specialised expertise, rigorous feasibility checks, and integrated simulation toolchains. Large language models (LLMs) have emerged as powerful assistants for scientific work, offering natural-language interfaces that can explain concepts, parse files, and generate code and documentation, thereby lowering the barrier to GSM interpretation and analysis workflow setup, accelerating hypothesis generation, and improving accessibility for non-experts. Despite this promise, limited evidence exists regarding LLMs’ domain knowledge in interpreting GSM and implementing the analysis tasks.

Here, we comprehensively assessed LLM capabilities in understanding and analyzing GSMs for metabolic engineering: (i) We systematically evaluated four core areas: domain knowledge, metabolic flux prediction, model reconstruction for pathway design, and metabolic flux optimization; (ii) We benchmarked four prominent LLMs (GPT, Gemini, Claude, and Deepseek-R1) and assessed their outputs using multi-model auto-evaluation approach based on standardized rubric-based scoring metrics, with independent cross-evaluations by all participating models; (iii) We identified recurrent failure modes and task- and model-specific limitations, which highlight areas for improvement; (iv) Based on our analysis, we have articulated best practices for integrating LLMs into GSM workflow for knowledge extraction and analysis implementation pipelines, facilitating their use for pathway and strain design. This work establishes an evidence-based baseline for LLM-enabled GSM analysis and informs the development of more reliable, accessible, and automation-ready computational workflows for metabolic engineering.

Comment

To browse or post comments, you must log in.Log in