Poster ID
P-35
Poster Title
An Agentic AI-Powered Genomic Data Dashboard in a Trusted Research Environment for Accelerated Data Exploration and Analysis
Authors
Brandon Xuan Ming Phua — Bioinformatics Institute (BII), Agency for Science, Technology and Research (A*STAR), Singapore
Zhihui Li — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Zheng Li — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Kar Seng Sim — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
AI Shan Lee — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Jun Xian Liew — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Sing-Wu Liou — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Zhenyuan Ng — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Ivy Ee Ling Quek — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Shan Ho Tan — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Arshia Naaz — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Shainan Hora — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Cheng-Shoong Chong — Bioinformatics Institute (BII), Agency for Science, Technology and Research (A*STAR), Singapore
Sebastian Maurer-Stroh — Bioinformatics Institute (BII), Agency for Science, Technology and Research (A*STAR), Singapore
Rajkumar Dorajoo — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Jianjun Liu — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Chih Chuan Shih — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Xing Yi Woo* — Bioinformatics Institute (BII), Agency for Science, Technology and Research (A*STAR), Singapore
*Corresponding: woo_xing_yi@a-star.edu.sg
Zhihui Li — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Zheng Li — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Kar Seng Sim — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
AI Shan Lee — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Jun Xian Liew — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Sing-Wu Liou — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Zhenyuan Ng — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Ivy Ee Ling Quek — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Shan Ho Tan — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Arshia Naaz — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Shainan Hora — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Cheng-Shoong Chong — Bioinformatics Institute (BII), Agency for Science, Technology and Research (A*STAR), Singapore
Sebastian Maurer-Stroh — Bioinformatics Institute (BII), Agency for Science, Technology and Research (A*STAR), Singapore
Rajkumar Dorajoo — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Jianjun Liu — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Chih Chuan Shih — Genome Institute of Singapore (GIS), Agency for Science, Technology and Research (A*STAR), Singapore
Xing Yi Woo* — Bioinformatics Institute (BII), Agency for Science, Technology and Research (A*STAR), Singapore
*Corresponding: woo_xing_yi@a-star.edu.sg
Abstract
National-scale whole-genome sequencing (WGS) projects generate extensive sample metadata, sequencing records, QC metrics, alignment files, VCFs, and otheroutputs. These data are increasingly accessed within trusted research environments (TREs) because of security data governance requirements (Stark et al., Nat Rev Genet 2025). However, restricted access, limited software, complex setup and approval processes (Kavianpour et al., J Med Internet Res 2022) often require skilled bioinformaticians to first gain familiarity of the TRE and build custom workflows even for basic data exploration. This increases computational time and cost while delaying the initial exploration needed to formulate research hypotheses.
We developed an interactive dashboard application that integrates sequencing-related information for approved genomic data sets within a TRE. Once access is granted, this system combines patient and laboratory metadata, sequencing records, QC reports, and VCF files in the approved workspace, converts them into structured layers for fast querying, and launch the dashboard for interactive user exploration and analysis.
Currently, this system provides six core capabilities: (1) interactive visualization and filtering of sequencing QC metrics; (2) browsing and interrogation of VCF data in a genome browser; (3) customizable visualization of variant summary statistics; (4) direct querying of variant data; (5) chat-based agent driven by a local LLM for further data exploration via natural language, and (6) AI-agent assisted initiation of downstream analytical workflows based on user preference. This system is tested on a cloud-based TRE system RAPTOR (Shih et al., iScience 2023) to ensure feasibility based on its security and governance framework using 2 Singapore cohorts (n~760) of germline WGS data with corresponding metadata, file manifest and QC metrics.
This agentic-AI interactive platform enables efficient exploration of large genomic datasets across cohorts without requiring users to install specialist tools or construct complex command-line queries, while being subjected to TRE governance approvals. It reduces the expertise, time, and computational overheads required to analyse governed genomic data while maintaining compliance with access and governance policies, supporting more secure, scalable, and accessible genomics research.
We developed an interactive dashboard application that integrates sequencing-related information for approved genomic data sets within a TRE. Once access is granted, this system combines patient and laboratory metadata, sequencing records, QC reports, and VCF files in the approved workspace, converts them into structured layers for fast querying, and launch the dashboard for interactive user exploration and analysis.
Currently, this system provides six core capabilities: (1) interactive visualization and filtering of sequencing QC metrics; (2) browsing and interrogation of VCF data in a genome browser; (3) customizable visualization of variant summary statistics; (4) direct querying of variant data; (5) chat-based agent driven by a local LLM for further data exploration via natural language, and (6) AI-agent assisted initiation of downstream analytical workflows based on user preference. This system is tested on a cloud-based TRE system RAPTOR (Shih et al., iScience 2023) to ensure feasibility based on its security and governance framework using 2 Singapore cohorts (n~760) of germline WGS data with corresponding metadata, file manifest and QC metrics.
This agentic-AI interactive platform enables efficient exploration of large genomic datasets across cohorts without requiring users to install specialist tools or construct complex command-line queries, while being subjected to TRE governance approvals. It reduces the expertise, time, and computational overheads required to analyse governed genomic data while maintaining compliance with access and governance policies, supporting more secure, scalable, and accessible genomics research.