bioRxiv Science⌕ Search

Biology subjects

Cozzini, S.

Publications and source records attributed to Cozzini, S..

3 recordsLinked to original sources

Scalable, fast and accurate differential gene expression testing from millions of cells of multiple patients

Since the development of DNA microarrays and later RNA bulk sequencing, testing with statistically independent samples has been the standard method for detecting genes with different transcription patterns. Single-cell assays challenge these assumptions because individual cells are statistically dependent, and all proposed methodologies present mathematical limitations or computational bottlenecks that prevent a seamless integration of data from many cells and patients simultaneously. In this work, we solve this crucial limitation by introducing a Bayesian framework that retrieves the independence structure at the level of individual patients, separating differences across individuals from actual transcriptional differences. Leveraging multi-GPU and variational inference, our approach excels across different experimental designs and scales to analyse over 10 million cells. This framework enables single-cell differential expression analysis that can finally integrate datasets from large clinical cohorts, atlas projects, or drug-response screens with thousands of samples and millions of cells.

bioinformatics↗

Unsupervised domain classification of AlphaFold2-predicted protein structures

AO_SCPLOWBSTRACTC_SCPLOWThe release of the AlphaFold database, which contains 214 million predicted protein structures, represents a major leap forward for proteomics and its applications. However, lack of comprehensive protein annotation limits its accessibility and usability. Here, we present DPCstruct, an unsupervised clustering algorithm designed to provide domain-level classification of protein structures. Using structural predictions from AlphaFold2 and comprehensive all-against-all local alignments from Foldseek, DPCstruct identifies and groups recurrent structural motifs into domain clusters. When applied to the Foldseek Cluster database, a representative set of proteins from the AlphaFoldDB, DPCstruct successfully recovers the majority of protein folds catalogued in established databases such as SCOP and CATH. Out of the 28,246 clusters identified by DPCstruct, 24% have no structural or sequence similarity to known protein families. Supported by a modular and efficient implementation, classifying 15 million entries in less than 48 hours, DPCstruct is well suited for large-scale proteomics and metagenomics applications. It also facilitates the rapid incorporation of updates from the latest structural prediction tools, ensuring that the classification remains up-to-date. The DPCstruct pipeline and associated database are freely available in a dedicated repository, enhancing the navigation of the AlphaFoldDB through domain annotations and enabling rapid classification of other protein datasets.

bioinformatics↗

Protein family annotation for the Unified Human Gastrointestinal Proteome by DPCfam clustering

Technological advances in massively parallel sequencing have led to an exponential growth in the number of known protein sequences. Much of this growth originates from metagenomic projects producing new sequences from environmental and clinical samples. The Unified Human Gastrointestinal Proteome (UHGP) catalogue is one of the most relevant metagenomic datasets with applications ranging from medicine to biology. However, the lack of sequence annotation impairs its usability. This work aims to produce a family classification of UHGP sequences to facilitate downstream structural and functional annotation. This is achieved through the release of the DPCfam-UHGP50 dataset containing 10,778 putative protein families generated using DPCfam clustering, an unsupervised pipeline grouping sequences into multi-domain architectures. DPCfam-UHGP50 considerably improves family coverage at protein and residue levels compared to the manually curated repository Pfam. It is our hope that DPCfam-UHGP50 will foster future discoveries in the field of metagenomics of the human gut by the release of a FAIR-compliant database easily accessible via a searchable web server and Zenodo repository.

bioinformatics↗