bioRxiv Science⌕ Search

Biology subjects

JAFARI, M.

Publications and source records attributed to JAFARI, M..

3 recordsLinked to original sources

CoPISA: Combinatorial Proteome Integral Solubility/Stability Alteration analysis

Combination therapies are widely used in acute myeloid leukemia (AML), but systematic datasets capturing proteome-wide responses to multi-drug perturbations remain limited. Here we present CoPISA (Combinatorial Proteome Integral Solubility/Stability Alteration), a quantitative proteomics assay designed to profile protein solubility and stability responses to single and combined drug treatments. The dataset includes two AML drug pairs (LY3009120-sapanisertib and ruxolitinib-ulixertinib) applied to four AML cell lines (MOLM-13, MOLM-16, SKM-1, and NOMO-1) under control, single-agent, and combination conditions in both lysate and intact-cell formats. Thermal solubility profiling coupled with TMT-based multiplexed LC-MS/MS generated 16 TMT16-plex experiments comprising 192 LC-MS/MS raw files, providing deep proteome coverage across treatments and biological contexts. The resource includes raw and processed proteomics data, detailed experimental metadata in Sample and Data Relationship Format (SDRF), and reproducible analysis scripts for reporter normalization, protein-level aggregation, statistical modeling, and classification of combinatorial response patterns. The experimental design enables identification of proteins responding uniquely to combination treatments as well as overlapping single-agent effects. Technical validation demonstrates reproducible quantification across multiplex experiments and assay formats. All data are publicly available through the PRIDE repository (PXD066812) together with analysis code, enabling independent reanalysis and method development. This dataset provides a benchmark resource for studying proteome responses to drug combinations, comparing lysate and intact-cell perturbation profiles, developing computational approaches for combinatorial target inference, and supporting training in computational proteomics.

bioinformatics↗

SOORENA: Self-lOOp containing or autoREgulatory Nodes in biological network Analysis

Autoregulatory mechanisms, in which proteins regulate their own activity or expression, are fundamental to biological networks but are challenging to identify systematically from literature. To address this gap, we present SOORENA (https://soorena.it.helsinki.fi/soorena/), a two-stage transformer model that predicts and classifies protein autoregulation in PubMed abstracts. SOORENA was trained on 1,332 experimentally validated abstracts and achieved 96.0 percent accuracy and 97.8 percent precision in stage one, with stage two achieving 95.5 percent accuracy and 96.2 percent macro-F1 across seven mechanistic classes. Applied to 3.34 million abstracts, SOORENA identified 85,145 publications containing autoregulatory mechanisms, yielding 97,657 protein-specific records. Integration with curated databases generated 100,065 comprehensive entries accessible via an interactive Shiny application. By systematically cataloging self-regulatory interactions, which often act as bottlenecks in dynamic network modeling, SOORENA provides a resource that supports mechanistic interpretation, model reduction, and predictive systems-level analyses. These results demonstrate that domain-specific language models can scale the discovery and curation of biologically essential self-regulatory mechanisms, bridging literature mining and systems biology.

bioinformatics↗

Missing Values Are Valuable: Shifting Focus from Amount to Form of Missing Data

Missing data is often treated as a nuisance, routinely imputed or excluded from statistical analyses, especially in nominal datasets where its structure cannot be easily modeled. However, the form of missingness itself can reveal hidden relationships, substructures, and biological or operational constraints within a dataset. In this study, we present a graph-theoretic approach that reinterprets missing values not as gaps to be filled, but as informative signals. By representing nominal variables as nodes and encoding observed or missing associations as edges, we construct both weighted and unweighted bipartite graphs to analyze modularity, nestedness, and projection-based similarities. This framework enables downstream clustering and structural characterization of nominal data based on the topology of observed and missing associations; edge prediction via multiple imputation strategies is included as an optional downstream analysis to evaluate how well inferred values preserve the structure identified in the non-missing data. Across a series of biological, ecological, and social case studies, including proteomics data, the BeatAML drug screening dataset, ecological pollination networks, and HR analytics, we demonstrate that the structure of missing values can be highly informative. These configurations often reflect meaningful constraints and latent substructures, providing signals that help distinguish between data missing at random and not at random. When analyzed with appropriate graph-based tools, these patterns can be leveraged to improve the structural understanding of data and provide complementary signals for downstream tasks such as clustering and similarity analysis. Our findings support a conceptual shift: missing values are not merely analytical obstacles but valuable sources of insight that, when properly modeled, can enrich our understanding of complex nominal systems across domains. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=107 SRC="FIGDIR/small/670516v2_ufig1.gif" ALT="Figure 1"> View larger version (24K): org.highwire.dtl.DTLVardef@99c5eaorg.highwire.dtl.DTLVardef@1909d8corg.highwire.dtl.DTLVardef@1578c93org.highwire.dtl.DTLVardef@ce2e90_HPS_FORMAT_FIGEXP M_FIG C_FIG Shiny app address https://ehsan-zangene.shinyapps.io/nimaa_app/

bioinformatics↗