Poster Details
Poster ID
P-07
Poster Title
Cross-level normalisation of literature-derived variants for interoperability with GA4GH VRS
Authors
Anais Mottaz
University of Applied Sciences and Arts of Western Switzerland, Geneva
Swiss Institute of Bioinformatics, Geneva, Switzerland

Worawich Phornsiricharoenphant
University of Zurich, Zurich, Switzerland
Swiss Institute of Bioinformatics, Zurich, Switzerland
National Center for Genetic Engineering and Biotechnology, Thailand
National Science and Technology Development Agency, Thailand

Emilie Pasche
University of Applied Sciences and Arts of Western Switzerland, Geneva
Swiss Institute of Bioinformatics, Geneva, Switzerland

Alexandre Flament
University of Applied Sciences and Arts of Western Switzerland, Geneva
Swiss Institute of Bioinformatics, Geneva, Switzerland

Luc Mottin
University of Applied Sciences and Arts of Western Switzerland, Geneva
Swiss Institute of Bioinformatics, Geneva, Switzerland

Michael Baudis
University of Zurich, Zurich, Switzerland
Swiss Institute of Bioinformatics, Zurich, Switzerland

Patrick Ruch
University of Applied Sciences and Arts of Western Switzerland, Geneva
Swiss Institute of Bioinformatics, Geneva, Switzerland
Abstract
Interpretation of genomic data depends on literature evidence, but published variants are difficult to integrate consistently. Descriptions may omit reference sequences, use heterogeneous nomenclatures, span protein, transcript and genomic coordinates, and represent variants absent from databases. Their integration therefore requires database-independent, cross-level normalisation and interoperable identifiers.

Developed within the Swiss Institute of Bioinformatics Literature Services (sibils.org) as the variant query-expansion engine behind the literature search engine Variomes (variomes.sibils.org), SynVar (synvar.sibils.org) is evaluated in a literature normalisation workflow. It recognizes substitutions, indels, duplications and protein frameshifts expressed in free text and structured variant representations across genomic, coding and protein levels, as well as dbSNP rsIDs and ClinGen CAIDs. It resolves implicit gene or chromosome references and normalises using UniProt, VariantValidator and Mutalyzer. Because mappings and protein back-translation may be non-unique, SynVar retains plausible representations through reference mapping and coordinate conversion without requiring database identifiers. Outputs include normalized HGVS and VCF-like representations, plus GA4GH VRS Alleles generated with the python reference implementation.

We assessed SynVar in two settings. First, we tested whether diverse variant descriptions across molecular levels and notation formats produced consistent linked representations. We evaluated 59 pathogenic ClinVar variants rewritten into up to 13 forms each (653 cases). SynVar interpreted and normalized 625 cases (95.7%) and matched ClinVar on 92-96% across output fields. Discrepancies reflected missing UniProt cross-references, ambiguous protein-frameshift back-translation and transcript selection. At the shared GRCh38 genomic level, all 625 cases produced the same GA4GH VRS computed identifiers as the VICC variation-normalizer.

Second, we evaluated the literature-annotation pipeline on 80 open-access PM3-Bench papers, from SIBiLS extraction of sentences, tables and genes, to SynVar variant recognition and normalisation. End-to-end recall was 58%. Most losses were due to forms not yet recognized by SynVar but readily addressable, absent gene annotations and dense variant tables.

Together, SIBiLS and SynVar have the potential to make literature-derived variant evidence interoperable with genomic data resources.
Digital Poster
View Poster
Close