bioRxiv Science⌕ Search

Biology subjects

Plancade, S.

Publications and source records attributed to Plancade, S..

2 recordsLinked to original sources

A combined test for feature selection on sparse metaproteomics data - alternative to missing value imputation

One of the difficulties encountered in the statistical analysis of metaproteomics data is the high proportion of missing values, which are usually treated by imputation. Nevertheless, imputation methods are based on restrictive assumptions regarding missingness mechanisms, namely "at random" or "not at random". To circumvent these limitations in the context of feature selection in a multi-class comparison, we propose a univariate selection method that combines a test of association between missingness and classes, and a test for difference of observed intensities between classes. This approach implicitly handles both missingness mechanisms. We performed a quantitative and qualitative comparison of our procedure with imputation-based feature selection methods on two experimental data sets, as well as simulated data with various scenarios regarding the missingness mechanisms and the nature of the difference of expression (differential intensity or differential missingness). Whereas we observed similar performances in terms of prediction on the experimental data set, the feature ranking and selection from various imputation-based methods were strongly divergent. We showed that the combined test reaches a compromise by correlating reasonably with other methods, and remains efficient in all simulated scenarios unlike imputation-based feature selection methods.

bioinformatics↗

A stochastic process modelling of maize phyllochron enables to characterize environmental and genetic effects.

The times between appearance of successive leaves or phyllochron characterize the vegetative development of annual plants. Hypothesis testing models, which enables to compare phyllochron between genetic groups or conditions, are usually based on regression of thermal time on the number of leaves, most of the time assuming a constant leaf appearance rate. However these models are both statistically biased and inappropriate in terms of modelling. We propose a stochastic process model in which the emergence of new leaves is considered as successive time-to-events, which provides a flexible and more accurate modelling as well as unbiased testing procedures. The model was applied on an original maize dataset collected in fields for three years on plants originating from two divergent selection experiments for flowering time conducted in two maize inbred lines. We showed that the main differences in phyllochron were not observed between selection populations (Early or Late), but rather between ancestral lines, years of experimentation, and leaf ranks. Our results highlight a strong departure from the assumption of a constant leaf appearance rate in one year that could be related to climate variations, even if the impact of each climatic variables individually was not clearly elucidated.

plant biology↗