bioRxiv ScienceSearch

Biology subjects

de Castro, G. M.

Publications and source records attributed to de Castro, G. M..

2 recordsLinked to original sources

CALANGO: an annotation-based, phylogeny-aware comparative genomics framework for exploring and interpreting complex genotypes and phenotypes

The increasing availability of genomic, annotation, evolutionary and phenotypic data for species contrasts with the lack of studies that adequately integrate these heterogeneous data sources to produce biologically meaningful knowledge. Here, we present CALANGO, a phylogeny-aware comparative genomics tool that uncovers functional molecular convergences and homologous regions associated with quantitative genotypes and phenotypes across species, enabling the fast discovery of novel statistically sound, biologically relevant phenotype-genotype associations. We demonstrate the usefulness of CALANGO in two case studies. The first one unveils potential causal links between prophage density and the pathogenicity phenotype in Escherichia coli, and confidently demonstrates how CALANGO supports the investigation of basic causal relationships by enabling a level of counterfactual investigation of observed associations in the data. As a second case study, we used our tool to search for homologous regions associated with a complex phenotypic trait in a major group of eukaryotes: the evolution of maximum height in angiosperms. We confidently identify a previously unknown association between maximum plant height and the expansion of the self-incompatibility system, a molecular mechanism that prevents inbreeding and increases genetic diversity. Taller species also have lower rates of molecular evolution due to their longer generation times, a critical concern for their long-term viability. The new mechanism we report could counterbalance this fact, and have far-reaching consequences for fields as diverse as conservation biology and agriculture. CALANGO is provided as a fully operational R package that can be freely installed from CRAN.

bioinformatics

Environmental DNA from a small sample of reservoir water can tell volumes about its biodiversity.

We evaluated the potential of metabarcoding in assessing the environmental DNA (eDNA) biodiversity profile in the water column of an hydroelectric power plant reservoir in southeast Brazil. Samples were obtained in three technical replicates at 1 km from the dam at 1, 13 and 25 m depths. For each minibarcodes -- COI, 12S and 16S -- 1.5 million paired-reads (150 base pairs) were sequenced. A total of 44 unique taxa were found. COI identified most of the taxa (34 taxa; 77.2 %) followed by 16S (14; 31.8 %) and 12S (10; 22.7 %). All minibarcodes identified fishes (13 taxa), however, COI detected other aquatic macro-invertebrates (18), algae (3) and amoebas (2). Richness was the same across the three depths (35 taxa), although, beta diversity suggested slightly divergent profiles. In just one location we identified 15 taxa never reported previously, 50% of the fish species identified in the last year of fishery monitoring and 13% of the species in biodiversity surveys performed from 2012 to 2021. Clustering into Amplicon Sequence Variants (ASV) showed that 12S and 16S are able to detect predominant haplotypes of fishes, suggesting they are suitable to study population genetics of this group. In this study we reviewed the species occurring within the Tres Irmaos reservoir according to previous conventional surveys and demonstrated that eDNA metabarcoding can be applied to monitor its biodiversity.

molecular biology