bioRxiv Science⌕ Search

Biology subjects

Hejblum, B. P.

Publications and source records attributed to Hejblum, B. P..

4 recordsLinked to original sources

Neglecting normalization impact in semi-synthetic RNA-seq data simulation generates artificial false positives

By reproducing differential expression analysis simulation results presented by Li et al, we identified a caveat in the data generation process. Data not truly generated under the null hypothesis led to incorrect comparisons of benchmark methods. We provide corrected simulation results that demonstrate the good performance of dearseq and argue against the superiority of the Wilcoxon rank-sum test as suggested by Li et al. Please see related Research article with DOI 10.1186/s13059-022-02648-4.

bioinformatics↗

High temporal resolution transcriptomic profiling delineates distinct patterns of interferon response following Covid-19 mRNA vaccination and SARS-CoV2 infection

Knowledge of the mechanisms underpinning the development of protective immunity conferred by mRNA vaccines is fragmentary. Here we investigated responses to COVID-19 mRNA vaccination via ultra-low-volume sampling and high-temporal-resolution transcriptome profiling (23 subjects across 22 timepoints, and with 117 COVID-19 patients used as comparators). There were marked differences in the timing and amplitude of the responses to the priming and booster doses. Notably, we identified two distinct interferon signatures. The first signature (A28/S1) was robustly induced both post-prime and post-boost and in both cases correlated with the subsequent development of antibody responses. In contrast, the second interferon signature (A28/S2) was robustly induced only post-boost, where it coincided with a transient inflammation peak. In COVID19 patients, a distinct phenotype dominated by A28/S2 was associated with longer duration of intensive care. In summary, high-temporal-resolution transcriptomic permitted the identification of post- vaccination phenotypes that are determinants of the course of COVID-19 disease.

immunology↗

Gene Set Analysis for time-to-event outcome with the Generalized Berk-Jones statistic

Gene set analysis evaluates the collective impact of groups of genes on an outcome of interest, such as disease occurrence. By incorporating biological knowledge through predefined gene sets, this approach enhances the interpretability of results and improves statistical power compared to gene-wise analyses. In the context of time-to-event data, existing methods are limited and fail to account for potentially strong correlations within gene sets. Given the strong performance of the Generalized Berk-Jones (GBJ) statistic, which effectively incorporates correlation within the test statistic, we adapted this method to the time-to-event framework using a Cox model. We then compared its performance with established methods, including the Wald test, global test, and global boost test. Our proposed method, sGBJ, shows an over-control of Type I error, leading to reduced statistical power compared to other methods in numerical studies. We further benchmarked these methods in two different real-world contexts: gliomas and breast cancer. The Wald test emerged as the most effective, identifying the largest number of significant pathways while maintaining appropriate control of Type I error in simulation settings. sGBJ closely followed demonstrating good performances, without a significant loss of statistical power in analyzing these two real-world biomedical datasets.

bioinformatics↗

Distribution-free complex hypothesis testing for single-cell RNA-seq differential expression analysis

SO_SCPLOWUMMARYC_SCPLOWState-of-the-art methods for single-cell RNA sequencing (scRNA-seq) Differential Expression Analysis (DEA) often rely on strong distributional assumptions that are difficult to verify in practice. Furthermore, while the increasing complexity of clinical and biological single-cell studies calls for greater tool versatility, the majority of existing methods only tackle the comparison between two conditions. We propose a novel, distribution-free, and flexible approach to DEA for single-cell RNA-seq data. This new method, called ccdf, tests the association of each gene expression with one or many variables of interest (that can be either continuous or discrete), while potentially adjusting for additional covariates. To test such complex hypotheses, ccdf uses a conditional independence test relying on the conditional cumulative distribution function, estimated through multiple regressions. We provide the asymptotic distribution of the ccdf test statistic as well as a permutation test (when the number of observed cells is not sufficiently large). ccdf substantially expands the possibilities for scRNA-seq DEA studies: it obtains good statistical performance in various simulation scenarios considering complex experimental designs (i.e. beyond the two condition comparison), while retaining competitive performance with state-of-the-art methods in a two-condition benchmark. We apply ccdf to a large publicly available scRNA-seq dataset of 84,140 SARS-CoV-2 reactive CD8+ T cells, in order to identify the diffentially expressed genes across 3 groups of COVID-19 severity (mild, hospitalized, and ICU) while accounting for seven different cellular subpopulations.

bioinformatics↗