bioRxiv ScienceSearch

Biology subjects

Brazeau, M. D.

Publications and source records attributed to Brazeau, M. D..

2 recordsLinked to original sources

Influence of different modes of morphological character correlation on phylogenetic tree inference

Phylogenetic analysis algorithms require the assumption of character independence - a condition generally acknowledged to be violated by morphological data. Correlation between characters can originate from intra-organismal features, shared phylogenetic history or forced by particular character-state coding schemes. Although the two first sources can be investigated by biologists a posteriori and the third one can be avoided a priori with good practices, phylogenetic software do not distinguish between any of them.\n\nIn this study, we propose a new metric of raw character difference as a proxy for character correlation. Using thorough simulations, we test the effect of increasing or decreasing character differences on tree topology. Overall, we found an expected positive effect of reducing character correlations on recovering the correct topology. However, this effect is less important for matrices with a small number of taxa (25 in our simulations) where reducing character correlation is not more effective than randomly drawing characters. Furthermore, in bigger matrices (350 characters), there is a strong effect of the inference method with Bayesian trees being consistently less affected by character correlation than maximum parsimony trees.\n\nThese results suggest that ignoring the problem of character correlation or independence can often impact topology in phylogenetic analysis. However, encouragingly, they also suggest that, unless correlation is actively maximised or minimised, probabilistic methods can easily accommodate for a random correlation between characters.

evolutionary biology

Morphological phylogenetic analysis with inapplicable data

Non-independence of characters is a real phenomenon in phylogenetic data matrices, even though phylogenetic reconstruction algorithms generally assume character independence. In morphological datasets, the problem results in characters that cannot be applied to certain terminal taxa, with this inapplicability treated as \"missing data\" in a popular method of character coding. However, this treatment is known to create spurious tree length estimates on certain topologies, potentially leading to erroneous results in phylogenetic searches. Here we present a single-character algorithm for ancestral states reconstruction in datasets that have been coded using reductive coding. The algorithm uses up to four traversals on a tree to resolve final ancestral states - which are required in full before a tree can be scored. The algorithm employs explicit criteria for the resolution of ambiguity in applicable/inapplicable dichotomies and the optimization of missing data. We score trees following a previously published procedure that minimizes homoplasy over all characters. Our analysis of published datasets shows that, compared to traditional methods, our new method identifies different trees as \"optimal\"; as such, correcting for inapplicable data may significantly alter the outcome of tree searches.

evolutionary biology