bioRxiv ScienceSearch

Biology subjects

Verissimo, A.

Publications and source records attributed to Verissimo, A..

4 recordsLinked to original sources

A haplotype-resolved draft genome of the European sardine (Sardina pilchardus)

BackgroundThe European sardine (Sardina pilchardus Walbaum, 1792) has a high cultural and economic importance throughout its distribution. Monitoring studies of sardine populations report an alarming decrease in stocks due to overfishing and environmental change, which has resulted in historically low captures along the Iberian Atlantic coast. Consequently, there is an urgent need to better understand the causal factors of this continuing decrease in the sardine stock. Important biological and ecological features such as levels of population diversity, structure, and migratory patterns can be addressed with the development and use of genomics resources.\n\nFindingsThe sardine genome of a single female individual was sequenced using Illumina HiSeq X Ten 10X Genomics linked-reads generating 113.8 Gb of data. Three draft genomes were assembled: two haploid genomes with a total size of 935 Mbp (N50 103Kb) each, and a consensus genome with a total size of 950 Mbp (N50 97Kb). The genome completeness assessment captured 84% of Actinopterygii Benchmarking Universal Single-Copy Orthologs. To obtain a more complete analysis, the transcriptomes of eleven tissues were sequenced and used to aid the functional annotation of the genome, resulting in 40 777 genes predicted. Variant calling on nearly half of the haplotype genome resulted in the identification of more than 2.3 million phased SNPs with heterozygous loci.\n\nConclusionsA draft genome was obtained with the 10X Genomics linked-reads technology, despite a high level of sequence repeats and heterozygosity that are expected genome characteristics of a wild sardine. The reference sardine genome and respective variant data are a cornerstone resource of ongoing population genomics studies to be integrated into future sardine stock assessment modelling to better manage this valuable resource.

genomics

Consensus outlier detection in survival analysis using the rank product test

Survival analysis is a well known technique in the medical field. The identification of individuals whose survival time is too short or to long given their profile, assumes great importance for the detection of new prognostic factors. The study of these outlying observations have gained increasing relevancy with the availability of high-throughput molecular and clinical data for large cohorts of patients. Several methods for outlier detection in survival data have been proposed, which include the analysis of the residuals, the measurement of the concordance c-index, and methods based on quantile regression for censored data. However, different results are obtained depending on the type of method used. In order to solve the disparity of results we proposed to apply the Rank Product test. A simulated dataset, and two clinical datasets were used to illustrate our proposed consensus outlier detection method, one from myeloma disease and the other from The Cancer Genome Atlas (TCGA) ovarian cancer. Finally, the Rank Product with multiple testing corrections was performed in order to identify which observations have the highest rank amongst the methods considered. Our results illustrate the potential of this consensus approach for the automated retrieval of outliers and also the identification of biomarkers associated with survival in large datasets.

bioinformatics

Sparse network-based regularization for the analysis of patientomics high-dimensional survival data

Data availability by modern sequencing technologies represents a major challenge in oncological survival analysis, as the increasing amount of molecular data hampers the generation of models that are both accurate and interpretable. To tackle this problem, this work evaluates the introduction of graph centrality measures in classical sparse survival models such as the elastic net.\n\nWe explore the use of network information as part of the regularization applied to the inverse problem, obtained both by external knowledge on the features evaluated and the data themselves. A sparse solution is obtained either promoting features that are isolated from the network or, alternatively, hubs, i.e., features that are highly connected within the network.\n\nWe show that introducing the degree information of the features when inferring survival models consistently improves the model predictive performance in breast invasive carcinoma (BRCA) transcriptomic TCGA data while enhancing model interpretability. Preliminary clinical validation is performed using the Cancer Hallmarks Analytics Tool API and the String database.\n\nThese case studies are included in the recently released glmSparseNet R package1, a flexible tool to explore the potential of sparse network-based regularizers in generalized linear models for the analysis of omics data.

bioinformatics

MassBlast: A workflow to accelerate RNA-seq and DNA database analysis

SummaryCurrent workflows for sequence analysis heavily depend on user input and manual curation. New specialized tools and methods are appearing all the time, but the actions required for a full analysis are disconnected and very time-consuming. The software we propose, MassBlast, combines BLAST+ and an automated workflow analysis to filter the results and significantly improve the annotation of multiple sequencing databases for exploring new biosynthetic pathways and new protein families, among other applications. MassBlast is fully configurable and reproducible.\n\nAvailability and ImplementationThe MassBlast package is written in Ruby. Source code and releases are freely available from Github (https://github.com/averissimo/mass-blast) for all major platforms (Linux, MS Windows and OS X) under the GPLv3 license.\n\nContactandre.verissimo@tecnico.ulisboa.pt

bioinformatics