bioRxiv Science⌕ Search

Biology subjects

Baxter, R.

Publications and source records attributed to Baxter, R..

2 recordsLinked to original sources

Low-coverage reduced representation sequencing reveals subtle within-island genetic structure in Aldabra giant tortoises

Aldabrachelys gigantea (Aldabra giant tortoise) is one of only two giant tortoise species left in the world and survives as a single wild population of over 100,000 individuals on Aldabra Atoll, Seychelles. Despite this large current population size, the species faces an uncertain future because of its extremely restricted distribution range and high vulnerability to the projected consequences of climate change. Captive-bred A. gigantea are increasingly used in rewilding programs across the region, where they are introduced to replace extinct giant tortoises in an attempt to functionally resurrect degraded island ecosystems. However, there has been little consideration of the current levels of genetic variation and differentiation within and among the islands on Aldabra. As previous microsatellite studies were inconclusive, we combined low-coverage and double digest restriction associated DNA (ddRAD) sequencing to analyze samples from 33 tortoises (11 from each main island). Using 5,426 variant sites within the tortoise genome, we detected patterns of population structure within two of the three studied islands, but no differentiation between the islands. These unexpected results highlight the importance of using genome-wide genetic markers to capture higher-resolution genetic structure to inform future management plans, even in a seemingly panmictic population. We show that low-coverage ddRAD sequencing provides an affordable alternative approach to conservation genomic projects of non-model species with large genomes.

ecology↗

Compositional Data Analysis using Kernels in Mass Cytometry Data

MotivationCell type abundance data arising from mass cytometry experiments are compositional in nature. Classical association tests do not apply to the compositional data due to their non-Euclidean nature. Existing methods for analysis of cell type abundance data suffer from several limitations for high-dimensional mass cytometry data, especially when the sample size is small. ResultsWe proposed a new multivariate statistical learning methodology, Compositional Data Analysis using Kernels (CODAK), based on the kernel distance covariance (KDC) framework to test the association of the cell type compositions with important predictors (categorical or continuous) such as disease status. CODAK scales well for high-dimensional data and provides satisfactory performance for small sample sizes (n < 25). We conducted simulation studies to compare the performance of the method with existing methods of analyzing cell type abundance data from mass cytometry studies. The method is also applied to a high-dimensional dataset containing different subgroups of populations including Systemic Lupus Erythematosus (SLE) patients and healthy control subjects. Availability and ImplementationCODAK is implemented using R. The codes and the data used in this manuscript are available on the web at http://github.com/GhoshLab/CODAK/. Supplementary informationSupplementary Materials.pdf.

bioinformatics↗