Poster ID
P-01
Poster Title
MedGen as a Shared Disease Terminology for GA4GH Standards
Authors
Adriana Malheiro, National Center for Biotechnology Information (NCBI), NLM, NIH
Megan S. Kane, National Center for Biotechnology Information (NCBI), NLM, NIH
Thilakam Venkatapathi, National Center for Biotechnology Information (NCBI), NLM, NIH
Douglas J Hoffman, National Center for Biotechnology Information (NCBI), NLM, NIH
Sahithi Punyasamudram, National Center for Biotechnology Information (NCBI), NLM, NIH
Brandi Kattman, National Center for Biotechnology Information (NCBI), NLM, NIH
Megan S. Kane, National Center for Biotechnology Information (NCBI), NLM, NIH
Thilakam Venkatapathi, National Center for Biotechnology Information (NCBI), NLM, NIH
Douglas J Hoffman, National Center for Biotechnology Information (NCBI), NLM, NIH
Sahithi Punyasamudram, National Center for Biotechnology Information (NCBI), NLM, NIH
Brandi Kattman, National Center for Biotechnology Information (NCBI), NLM, NIH
Abstract
Successful implementation of genomic medicine requires interoperable systems that connect patient phenotypes, test discovery, variant classification, and clinical interpretation. The clinical and research communities use overlapping terminologies, including HPO, OMIM®, Mondo, ORDO, SNOMED CT, and UMLS. MedGen, an NCBI resource for human medical genetics, integrates terms from these sources and from GeneReviews®, GTR and ClinVar submissions into records with Concept Unique Identifiers (CUIs), preferred names, synonyms, source identifiers, inheritance patterns, clinical features, gene relationships, hierarchies, definitions. These records have links to tests, variants, literature, authoritative sources, and clinical practice guidelines. MedGen can incorporate new concepts to support expert group specific terms, newly described phenotypes in literature, laboratories and data submitters to ClinVar and GTR, thus allowing the vocabulary to grow alongside the community's evolving knowledge. MedGen curators review conflicting mappings, split or merge concepts, engage with community curation efforts, and report discrepancies back to improve data harmonization among sources. MedGen is free and accessible via E-utilities API, FTP, and web (www.ncbi.nlm.nih.gov/medgen). This combination of breadth, cross-referencing, and curation makes MedGen well suited as a shared disease vocabulary across GA4GH workstreams.
In Genomic Knowledge Standards, MedGen can be used to anchor disease context for variant assertions, gene-disease evidence, and variant interpretation. In Clinical & Phenotype Data Capture, MedGen CUIs can complement HPO terms in Phenopackets, reports, and family history while preserving original text. In Discovery, MedGen can expand Beacon and Data Connect queries across condition names, CUIs, genes, phenotypes. In Federated Analysis, MedGen-coded conditions can support reproducible cohort definitions. MedGen's strong coverage of medical specialties, pharmacogenomics and rare diseases fills gaps relevant to multiple Driver Projects.
This work was supported by the National Center for Biotechnology Information of the National Library of Medicine (NLM), National Institutes of Health (NIH). The contributions of the NIH author(s) are considered Works of the United States Government. The findings and conclusions presented in this paper are those of the author(s) and do not necessarily reflect the views of the NIH or the U.S. Department of Health and Human Services.
In Genomic Knowledge Standards, MedGen can be used to anchor disease context for variant assertions, gene-disease evidence, and variant interpretation. In Clinical & Phenotype Data Capture, MedGen CUIs can complement HPO terms in Phenopackets, reports, and family history while preserving original text. In Discovery, MedGen can expand Beacon and Data Connect queries across condition names, CUIs, genes, phenotypes. In Federated Analysis, MedGen-coded conditions can support reproducible cohort definitions. MedGen's strong coverage of medical specialties, pharmacogenomics and rare diseases fills gaps relevant to multiple Driver Projects.
This work was supported by the National Center for Biotechnology Information of the National Library of Medicine (NLM), National Institutes of Health (NIH). The contributions of the NIH author(s) are considered Works of the United States Government. The findings and conclusions presented in this paper are those of the author(s) and do not necessarily reflect the views of the NIH or the U.S. Department of Health and Human Services.