Poster Details
Poster ID
P-13
Poster Title
The Rare Disease Data Commons: A governed, harmonized data and AI benchmarking platform for rare disease diagnostic methods
Authors
Gaia Andreoletti¹, Jineta Banerjee¹, Danielle Carnival², Luca Foschini¹, Rachel Galbraith³,
Benjamin Gyori⁴, Samuel Minot³, Milen Nikolov¹, Anthony Peña¹, Sasha Scott¹, Christine Suver¹,
Susheel Varma¹, Mark Weston ⁵ , Elizabeth H. Williams¹, Robert Allaway¹*

Affiliations:
1 Sage Bionetworks, Seattle, WA
2 Undiagnosed Diseases Network Foundation, Washington, DC
3 Cirro Bio, Seattle, WA
4 Northeastern University, Boston, MA
5 Netrias, Annapolis, MD
*with exception of the presenting author, all authors are listed in alphabetical order
Abstract
Rare disease research is slowed by data that are locked in institutional silos, inconsistently
annotated, and hard to access under defensive governance. More than 10,000 conditions affect
an estimated 350 million people, yet about half of patients are undiagnosed or misdiagnosed,
and diagnosis takes six years on average. To address this, we are building a Rare Disease
Data Commons (RDDC). RDDC will ingest, harmonize, and serve multimodal rare disease data
so diagnostic algorithms can be developed and evaluated against shared, well-governed
cohorts.

RDDC ingests multimodal rare disease data spanning longitudinal clinical records and remotely
captured, patient-generated measures, drawn from multiple upstream data collection efforts. As
the platform layer, it inherits the consent terms and data use agreements set by the providers,
encoding them as machine-readable access rules. The Commons manages ingestion, PHI
removal, and AI-assisted curation, integrating imaging, omics, patient-captured data, and clinical
records into analysis-ready collections. It implements GA4GH standards such as DUO for
computable governance, DRS for data access, and Phenopackets for phenotype capture.
Three features distinguish RDDC. First, governance is computable: DUO-encoded consent and
data-use restrictions operate within a Five Safes framework with element-level access controls
and provenance tracking, designed to support accountability to contributors, including the aim of
reporting back to communities on how their data are reused. Second, harmonization is
computational rather than manual: AI-assisted methods normalize heterogeneous data and
ground it to shared ontologies, enabling interoperability and linkage to external knowledge
resources. Third, benchmarking is shared infrastructure: through a model-to-data design,
algorithms are evaluated continuously against governed real or synthetic cohorts without
moving sensitive data. Where data cannot be centralized, RDDC supports federated integration
and interoperates with other trusted research environments rather than becoming another silo.
We propose RDDC as a candidate Driver Project and seek collaborations with other rare
disease databases and collections. Partners can contribute data for RDDC to host on their
behalf, or work with us on shared development across international platforms. The design
operates independently of its originating program, aiming to become durable community
infrastructure.

Funding: This research was, in part, funded by an award to RJA, JB, MN from the Advanced
Research Projects Agency for Health (ARPA-H). The views and conclusions contained in this
document are those of the authors and should not be interpreted as representing the official
policies, either expressed or implied, of the United States Government.
Close