Presentation Information
[P01-005]SynBioDB: Synthetic Biology Database
○Chinari Pawan Kumar Patro1,2, Wen Shan Yew1,2, Chueh Loo Poh1,2 (1. National University of Singapore (Singapore), 2. National Centre for Engineering Biology (Singapore))
Keywords:
Database,Artificial Intelligence,Synthetic Biology,Protein language models
A one-stop resource for Synthetic Biology data is highly essential together with unified data formats and data standards that are widely accepted for Synthetic Biology. To address this need, we have developed a database called “SynBioDB” – Synthetic Biology Database - at National Centre for Engineering Biology (NCEB), Singapore. SynBioDB is designed to enable efficient management of diverse biological data generated in the field of Synthetic Biology and, importantly, to support collaboration among different labs and institutions through standardized and unified data formats, governed by defined data standards (e.g., ISO, STRENDA, MIFlowCyt among others) and FAIR data principles.
The database supports raw and processed data of biological components and biological measurements (including metadata) pertaining to Synthetic Biology to support effective development of machine learning/AI models for engineering biology. It will be a one-stop hub for synthetic biology data comprising various biological components (e.g., bio-parts, enzymes, pathways and strains) and encompassing standard biological measurements including microplate reader data, flow cytometry, HPLC, mass spectrometry (LCMS and GCMS) data and fermentation data among others. The data in the database are stored with specific access permissions which allows the data to be restricted to specific lab/group of interest for data security.
In this talk, I will present about the various features of the database notably AI augmentation using protein language models (PLMs) and how it enables in finding further insights, user-friendly data upload, data download and automated data analysis; I will also discuss the data standards and the data formats that facilitate data storage in the database followed by analyses and generating plots. The database is open-source and is under GPL v3 license. Currently, the database is in its beta phase, and we welcome co-development. SynBioDB will eventually enable Machine Learning/AI-powered Design-Build-Test-Learn (DBTL) cycle which is at the core of Synthetic Biology. Together with the database, I will also present about other tools like Enzyme Function Initiative (EFI) tools that we are setting up at the NCEB, Singapore that is used for obtaining substantial new enzyme candidates that have not been tested further using sequence similarity networks (SSNs); and this facilitates sequence data and associated metadata for these new enzymes to be systematically stored in SynBioDB which can be further used for machine learning approaches completing the Design-Build-Test-Learn (DBTL) cycle.
The database supports raw and processed data of biological components and biological measurements (including metadata) pertaining to Synthetic Biology to support effective development of machine learning/AI models for engineering biology. It will be a one-stop hub for synthetic biology data comprising various biological components (e.g., bio-parts, enzymes, pathways and strains) and encompassing standard biological measurements including microplate reader data, flow cytometry, HPLC, mass spectrometry (LCMS and GCMS) data and fermentation data among others. The data in the database are stored with specific access permissions which allows the data to be restricted to specific lab/group of interest for data security.
In this talk, I will present about the various features of the database notably AI augmentation using protein language models (PLMs) and how it enables in finding further insights, user-friendly data upload, data download and automated data analysis; I will also discuss the data standards and the data formats that facilitate data storage in the database followed by analyses and generating plots. The database is open-source and is under GPL v3 license. Currently, the database is in its beta phase, and we welcome co-development. SynBioDB will eventually enable Machine Learning/AI-powered Design-Build-Test-Learn (DBTL) cycle which is at the core of Synthetic Biology. Together with the database, I will also present about other tools like Enzyme Function Initiative (EFI) tools that we are setting up at the NCEB, Singapore that is used for obtaining substantial new enzyme candidates that have not been tested further using sequence similarity networks (SSNs); and this facilitates sequence data and associated metadata for these new enzymes to be systematically stored in SynBioDB which can be further used for machine learning approaches completing the Design-Build-Test-Learn (DBTL) cycle.
Comment
To browse or post comments, you must log in.Log in
