Poster ID
P-08
Poster Title
SPDI v2: Compact VRS-Aligned Serialization for Structural Variant Identifiers
Authors
Lon Phan/NCBI
Tim Hefferon/NCBI
Alex Wagner/Nationwide Children’s Hospital, Ohio State University
College of Medicine
Larry Babb/Broad Institute
Tim Hefferon/NCBI
Alex Wagner/Nationwide Children’s Hospital, Ohio State University
College of Medicine
Larry Babb/Broad Institute
Abstract
Structural variants are challenging to represent consistently because they can involve large sequence changes, symbolic alleles, repeat-length differences, uncertain breakpoints, or novel adjacencies between distant genomic loci. GA4GH Variation Representation Specification (VRS) provides a computable object model for variation, and NCBI’s Sequence Position Deletion Insertion (SPDI) notation provides a compact, sequence-based syntax used for variant identification and normalization. This poster presents SPDI v2 as a compact, VRS-aligned string serialization for structural variants, with an emphasis on practical structural variant identifiers for use in genomic resources.
SPDI v2 extends the core sequence:position:deletion:insertion structure to support structural variant classes through compact representation based on their inherent properties. Repeat-associated variants can be represented using compact integer-based expressions rather than expanded sequence strings. Structural alleles can encode deletions, insertions, duplications, inversions, copy-number changes, and associated length or copy-number attributes. Adjacency-based representations can also describe junctions between distinct sequence locations, including orientation where required, enabling detailed descriptions of translocations and complex variants.
The design aims to produce compact strings suitable for data exchange, indexing, and computed identifiers while maintaining alignment and bidirectional interoperability with VRS concepts such as Allele and Adjacency. We also address key implementation challenges, including adjacency normalization, complex event representation, canonical ordering of multi-junction structures, and the distinction between lightweight identifier strings and richer VRS object models. SPDI v2 provides a pragmatic bridge between GA4GH variation models and NCBI’s operational needs for stable, concise structural variant identifiers.
SPDI v2 extends the core sequence:position:deletion:insertion structure to support structural variant classes through compact representation based on their inherent properties. Repeat-associated variants can be represented using compact integer-based expressions rather than expanded sequence strings. Structural alleles can encode deletions, insertions, duplications, inversions, copy-number changes, and associated length or copy-number attributes. Adjacency-based representations can also describe junctions between distinct sequence locations, including orientation where required, enabling detailed descriptions of translocations and complex variants.
The design aims to produce compact strings suitable for data exchange, indexing, and computed identifiers while maintaining alignment and bidirectional interoperability with VRS concepts such as Allele and Adjacency. We also address key implementation challenges, including adjacency normalization, complex event representation, canonical ordering of multi-junction structures, and the distinction between lightweight identifier strings and richer VRS object models. SPDI v2 provides a pragmatic bridge between GA4GH variation models and NCBI’s operational needs for stable, concise structural variant identifiers.