Genomic data is only as findable and reusable as the metadata describing how it was produced, and aligning that metadata with the demands of AI is central to GA4GH's mission. This session walks through the early data lifecycle by area of metadata capture, mapping GA4GH standards to AI-readiness layers: study, dataset, and cohort context (Beacon, Data Connect, provenance work), participant, phenotype and biosample (Phenopackets), experiment (Experiments Metadata Checklist), and data quality (Whole Genome Sequencing Quality Control Standards).
It also explores frameworks like FAIRSCAPE (using RO-Crate, PROV-O/EVI) to capture the provenance, ethics, and computability needed to make data genuinely AI-ready. The session brings people together to agree on practical metadata requirements, and prioritises engagement with local and broader Asian genomics communities, so regional needs and use cases shape the recommendations.
Delivered as an interactive workshop and community consultation, it spans the Clinical & Phenotypic Data Capture (Clin/Pheno) Work Stream and Discovery Work Streams, with Federated Analysis Work Stream involvement.
More details available here: https://docs.google.com/document/d/1JNnocRPvvV9-4dQHqueHnKbh7CQI-uaBxmjh6-Fwi3k/