Poster ID
P-42
Poster Title
Design and Implementation of a Sovereign Bioinformatics Infrastructure for National Pharmacogenomics Initiatives in Indonesia
Authors
Syahid Naufal Ramadhan [1], Stefiani Emasurya [1], Jessica Audrienna [1], Pamela Gan [1], Satrio Wibowo [2], Annisa Muthiah Sukirman [2], BGSI Ecosystem [2]
[1] PT Nalagenetik Riset Indonesia, Jakarta, Indonesia
[2] Biomedical and Genome Science Initiative (BGSI), Jakarta, Indonesia
[1] PT Nalagenetik Riset Indonesia, Jakarta, Indonesia
[2] Biomedical and Genome Science Initiative (BGSI), Jakarta, Indonesia
Abstract
National genomics initiatives increasingly incorporate specialized tertiary analyses, such as pharmacogenomics (PGx) interpretation, which demand niche expertise including curated allele-function knowledge and continuously updated clinical guidelines, that can be costly to build and maintain. This raises a practical question: how can data custodians access specialized bioinformatics capability without relinquishing control over sensitive genomic data? We present a federated analysis collaboration model built on GA4GH's data visiting principle, developed between a startup and a national precision medicine initiative in a lower middle income country, in which analytical tooling is delivered as portable, containerized software that executes entirely within the data holder's own cloud environment. Nextflow is used to orchestrate the pipeline, with each component, including third party tools and custom scripts and an embedded knowledgebase of CPIC, DPWG, and regulatory (FDA/EMA/PMDA) annotations, being packaged within docker containers for reproducible analysis. Processes are executed as AWS Batch tasks within the national initiative's own AWS account. Terraform templates let the data holder inspect and modify the infrastructure, while running within their own AWS account gives them full native visibility into execution logs and resource usage via standard AWS tooling (e.g., CloudTrail, CloudWatch). All CRAM/VCF processing, diplotype calling, and interpretation occur inside the national initiative's own AWS account; genomic data never leaves that boundary, and the knowledgebase is queried locally within the container, eliminating external API calls. This decoupling of "who controls the data" from "who develops the analysis" allowed specialized expertise to operate at population scale with no additional data transfer or custody arrangements. The model is mutually beneficial: the national initiative retains full custody of its data, while the startup avoids the cost and liability of directly handling sensitive data, allowing it to scale across jurisdictions. This work demonstrates a generalizable pattern for federated, data-visiting analysis, offering other national genomics programs a practical blueprint for public-private collaboration.