bioRxiv Science⌕ Search

Biology subjects

Pais, L.

Publications and source records attributed to Pais, L..

3 recordsLinked to original sources

Building an Interoperable Rare Disease Multi-omic Resource: The GREGoR Data Model and Dataset

Rare disease research and diagnosis rely on the integration of genomic and phenotypic data generated across diverse clinical sites; however, the absence of widely adopted standards for representing genomic data and associated metadata has limited data interoperability, reuse, and cross-study analysis. The Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) Consortium was established to investigate challenging rare disease cases and evaluate emerging multi-omic technologies for clinical translation. To support coordinated data integration across distributed research sites, we developed a common Consortium Data Model in partnership with domain experts to standardize the capture of participant-, family-, phenotype- and assay-level metadata, with a particular emphasis on using a modular architecture to support linking of multiple data versions from multiple omic technologies to a single individual and attribution of a genetic finding to the specific technology used for its initial discovery. Adoption of the GREGoR Data Model has enabled continued generation and public release of a harmonized, analysis-ready Consortium Dataset. The most recent release includes phenotypic, family and multi-omic data from 12,292 participants in 5,029 families. Other rare disease data sharing efforts are beginning to adopt this data model which will facilitate cross consortium analyses and empower rare disease research. This work demonstrates that a collaborative, flexible, and scalable data model can enable large-scale rare disease research, facilitate cross-center data harmonization, and enable data interoperability.

genomics↗

Evidence Aggregator: AI reasoning applied to rare disease diagnostics

Variant assessment of rare disease diagnostics depends on using domain knowledge in the time- consuming process of retrieving, reviewing, and synthesizing clinical and technical information. To address these challenges, we developed the Evidence Aggregator (EvAgg), an open-source, generative-AI-based tool designed for rare disease diagnosis that systematically extracts relevant information from the scientific literature for any human gene. EvAgg provides a thorough and current summary of observed genetic variants and their associated clinical features, enabling rapid synthesis of evidence concerning gene-disease relationships. We constructed an expert-curated dataset and evaluated EvAggs performance. EvAgg achieves 92% recall in identifying relevant papers, 96% recall in detecting instances of genetic variation within those papers, and [~]80% accuracy in extracting individual case and variant-level content (e.g. zygosity, inheritance, variant type, and phenotype). Further, EvAgg complemented the process of manual literature review by identifying substantial additional relevant information. When tested with analysts in rare disease case analysis, EvAgg reduced review time by 34% (p-value < 0.002) and increased the number of papers, variants, and cases evaluated per unit time. These savings have the potential to reduce diagnostic latency and increase solve rates for challenging rare disease cases.

genomics↗

Gene identification for ocular congenital cranial motor neuron disorders using human sequencing, zebrafish screening, and protein binding microarrays

PurposeTo functionally evaluate novel human sequence-derived candidate genes and variants for unsolved ocular congenital cranial dysinnervation disorders (oCCDDs). MethodsThrough exome and genome sequencing of a genetically unsolved human oCCDD cohort, we previously identified variants in 80 strong candidate genes. Here, we further prioritized a subset of these (43 human genes, 57 zebrafish genes) using a G0 CRISPR/Cas9-based knockout assay in zebrafish and generated F2 germline mutants for seventeen. We tested the functionality of variants of uncertain significance in known and novel candidate transcription factor-encoding genes through protein binding microarrays. ResultsWe first demonstrated the feasibility of the G0 screen by targeting known oCCDD genes phox2a and mafba. 70-90% of gene-targeted G0 zebrafish embryos recapitulated germline homozygous null-equivalent phenotypes. Using this approach, we then identified three novel candidate oCCDD genes (SEMA3F, OLIG2, and FRMD4B) with putative contributions to human and zebrafish cranial motor development. In addition, protein binding microarrays demonstrated reduced or abolished DNA binding of human variants of uncertain significance in known and novel sequence-derived transcription factors PHOX2A (p.(Trp137Cys)), MAFB (p.(Glu223Lys)), and OLIG2 (p.(Arg156Leu)). ConclusionsThis study nominates three strong novel candidate oCCDD genes (SEMA3F, OLIG2, and FRMD4B) and supports the functionality and putative pathogenicity of transcription factor candidate variants PHOX2A p.(Trp137Cys), MAFB p.(Glu223Lys), and OLIG2 p.(Arg156Leu). Our findings support that G0 loss-of-function screening in zebrafish can be coupled with human sequence analysis and protein binding microarrays to aid in prioritizing oCCDD candidate genes/variants.

neuroscience↗