bioRxiv Science⌕ Search

Biology subjects

Shilin, A.

Publications and source records attributed to Shilin, A..

3 recordsLinked to original sources

Predicted Effector Gene Aggregation, Standards and Unified Schema (PEGASUS): A Community Framework for Effector Gene Reporting

Genome-wide association studies (GWAS) increasingly report predicted effector genes (PEGs) - genes hypothesised to mediate the biological effects of associated variants. These function as key outputs for advancing variant-to-function research, mechanistic understanding, and therapeutic discovery. However, the rapid growth of PEG lists has not been matched by standards for organising, annotating, and reporting these predictions. As shown by recent landscape analyses, PEG lists vary widely in methodology, evidence definition, nomenclature, provenance tracking, and data structure, limiting interoperability, benchmarking, reuse, and adherence to FAIR principles. To address this gap, we convened an international multi-stakeholder community comprising method developers, data generators, resource maintainers, curators, funders, journal editors, and downstream users. Through a 2024 workshop and a 2025 working group series, we developed the Predicted Effector Gene Aggregation, Standards and Unified Schema (PEGASUS) framework. PEGASUS specifies (i) a metadata standard to capture provenance, trait and GWAS descriptors, evidence sources, and integration methods; (ii) a structured evidence matrix reporting all genes and all evidence underpinning prioritisation at each locus; and (iii) a concise PEG list that summarises author-prioritised genes linked transparently to underlying evidence. The framework balances transparency, machine readability, burden on submitters, and alignment with existing community standards. PEGASUS provides the first community-developed schema for reporting predicted effector genes and their supporting evidence. Adoption of this framework by authors will improve the comparability, reproducibility, and reusability of PEG outputs across studies, facilitating more robust biological inference, enabling cross-resource comparison of gene-prioritisation methods to support community benchmarking, and integration into downstream resources and analytical pipelines. PEGASUS-compliant data can be shared via the PEG Data Registry platform (https://kpndataregistry.org/peg), promoting re-use and establishing the basis for future integration with publicly shared GWAS data.

genomics↗

PanKbase Integrated Single-Cell Map: A Comprehensive Atlas of Human Pancreatic Islets

Single-cell RNA sequencing (scRNA-seq) of human pancreatic islet tissue is a powerful tool for investigating type 1 diabetes (T1D). However, individual datasets are limited in size and fragmented across donors, laboratories, and experimental conditions. To address this, we constructed a comprehensive, integrated scRNA-seq atlas of isolated human pancreatic islets by collating publicly available data generated from tissue provided by resources including the Human Pancreas Analysis Program, the Integrated Islet Distribution Program, and Prodo Labs. Systematic quality controls were implemented to select high-quality samples, reads, and cells. During integration, we accounted for important variables such as age, sex, body mass index, origin study, treatments, islet distribution resources, and sequencing chemistry. Our single-cell atlas comprises 191 high-quality samples from 140 donors (59 female, 81 male) across five phenotypic groups: no diabetes (controls, n=69), autoantibody positivity without diabetes (n=12), pre-diabetes (n=11), T1D (n=12), and type 2 diabetes (T2D) (n=36). In total, the atlas contains 448,935 cells, capturing 13 distinct populations, including alpha cells (43.3%) and beta cells (26.8%), as well as groups such as immune cells (0.6%). Publicly available at www.pankbase.org, this atlas provides a platform for hypothesis-driven investigation of diabetes pathophysiology and, given rigorous quality control, is well-suited for downstream machine-learning applications. Article HighlightsO_LICurrent scRNA-seq datasets of pancreatic islet tissue are limited in size and scattered across donors, laboratories, and experimental conditions, underscoring the need for a consolidated resource. C_LIO_LIWe harmonized datasets from multiple sources to build a comprehensive single-cell map of isolated human pancreatic islets. C_LIO_LIOur atlas captures 448,935 cells from 191 high-quality samples across 140 donors and multiple phenotypic groups, identifying 13 distinct cell populations. C_LIO_LIAvailable at www.pankbase.org, the atlas provides a scalable, rigorously curated platform to support hypothesis-driven diabetes research and can enable a broad range of downstream computational applications. C_LI

genomics↗

The Common Fund Data Ecosystem (CFDE)

The NIH Common Fund Data Ecosystem (CFDE) integrates data resources from 18 NIH Common Fund programs for discovery and integrative analysis. These programs generate valuable but heterogeneous datasets that can be difficult to discover, access, and reuse. CFDE aims to provide a collaborative, community-built infrastructure that links and enriches Common Fund programs. We describe the evolution, structure, and core technologies of CFDE, including practical approaches that support submission, integration, visualization, and public release of multimodal data. Training programs and workforce initiatives lower barriers to adoption. CFDE has devised solutions to critical issues facing cross-program initiatives, including data scale and heterogeneity, dataset integration, and long-term sustainability. We demonstrate the utility of linking Common Fund resources through integrative tools and cross-dataset queries to yield insights that would otherwise be infeasible. Collectively, CFDE shows that a standards-driven, federated approach enhances and unifies cross-disciplinary resources, fostering collaboration and data-driven discovery.

scientific communication and education↗