bioRxiv Science⌕ Search

Biology subjects

Saitta, A.

Publications and source records attributed to Saitta, A..

3 recordsLinked to original sources

A deep learning approach for improved detection of homologous recombination deficiency from shallow genomic profiles

Homologous Recombination Deficiency (HRD) is a predictive biomarker of poly-ADP ribose polymerase 1 inhibitors (PARPi) response. Most HRD detection methods are based on genome wide enumeration of scarring events and require deep genome sequence profiles (> 30x). The cost and workflow-specific biases introduced by these genome profiling methods currently limits clinical adoption of HRD testing. We introduce the Genomic Integrity Index (GII), a Convolutional Neuronal Network, that leverages features from low pass (1x) Whole Genome Sequencing data to distinguish HRD positive and negative samples. In a cohort of 230 ovarian and breast cancer, we found GII supports accurate stratification of samples yielding results that are highly concordant with state-of-the-art HRD detection methods (0.865<AUC<0.996) which require 50x deeper coverage. We conclude that the deep learning framework supporting GII allows accurate detection of HRD from shallow genome profiles, reducing biases and data generation costs making it uniquely suited for clinical applications.

cancer biology↗

Towards understanding diversity, endemicity and global change vulnerability of soil fungi

Fungi play pivotal roles in ecosystem functioning, but little is known about their global patterns of diversity, endemicity, vulnerability to global change drivers and conservation priority areas. We applied the high-resolution PacBio sequencing technique to identify fungi based on a long DNA marker that revealed a high proportion of hitherto unknown fungal taxa. We used a Global Soil Mycobiome consortium dataset to test relative performance of various sequencing depth standardization methods (calculation of residuals, exclusion of singletons, traditional and SRS rarefaction, use of Shannon index of diversity) to find optimal protocols for statistical analyses. Altogether, we used six global surveys to infer these patterns for soil-inhabiting fungi and their functional groups. We found that residuals of log-transformed richness (including singletons) against log-transformed sequencing depth yields significantly better model estimates compared with most other standardization methods. With respect to global patterns, fungal functional groups differed in the patterns of diversity, endemicity and vulnerability to main global change predictors. Unlike -diversity, endemicity and global-change vulnerability of fungi and most functional groups were greatest in the tropics. Fungi are vulnerable mostly to drought, heat, and land cover change. Fungal conservation areas of highest priority include wetlands and moist tropical ecosystems.

ecology↗

Guidelines for accurate genotyping of SARS-CoV-2 using amplicon-based sequencing of clinical samples

BackgroundSARS-CoV-2 genotyping has been instrumental to monitor virus evolution and transmission during the pandemic. The reliability of the information extracted from the genotyping efforts depends on a number of aspects, including the quality of the input material, applied technology and potential laboratory-specific biases. These variables must be monitored to ensure genotype reliability. The current lack of guidelines for SARS-CoV-2 genotyping leads to inclusion of error-containing genome sequences in studies of viral spread and evolution. ResultsWe used clinical samples and synthetic viral genomes to evaluate the impact of experimental factors, including viral load and sequencing depth, on correct sequence determination using an amplicon-based approach. We found that at least 1000 viral genomes are necessary to confidently detect variants in the genome at frequencies of 10% or higher. The broad applicability of our recommendations was validated in >200 clinical samples from six independent laboratories. The genotypes of clinical isolates with viral load above the recommended threshold cluster by sampling location and period. Our analysis also supports the rise in frequency of 20A.EU1 and 20A.EU2, two recently reported European strains whose dissemination was favoured by travelling during the summer 2020. ConclusionsWe present much-needed recommendations for reliable determination of SARS-CoV-2 genome sequence and demonstrate their broad applicability in a large cohort of clinical samples.

genomics↗