bioRxiv Science⌕ Search

bioRxiv · 10.1101/2025.05.09.653146

FEMA-Long: Modeling unstructured covariances for discovery of time-dependent effects in large-scale longitudinal datasets

Abstract

While linear mixed-effects (LME) models are common for analyzing longitudinal data, most users rely on random intercepts or simple stationary covariance, due to unavailability of computationally tractable solutions. Here, we extend the Fast and Efficient Mixed-Effects Algorithm (FEMA) and present FEMA-Long, a computationally tractable approach to flexibly modeling longitudinal covariance suitable for high-dimensional data. FEMA-Long can: i) model unstructured covariance, ii) model covariates as smooth functions using splines, iii) discover time-dependent effects of covariates with spline interactions, and iv) use these flexible longitudinal modeling strategies to perform longitudinal genome-wide association studies and discover time-dependent genetic effects, in a computationally scalable manner, suitable for high-dimensional data. Through extensive simulations, we show that estimates from FEMA-Long are accurate, while being up to several thousand times faster and with minimal carbon footprint. To show the utility of FEMA-Long for discovering novel biological signal, using data from the Norwegian Mother, Father and Child Cohort Study (MoBa), we performed a longitudinal genome-wide association study with non-linear SNP-by-time interaction on length, weight, and BMI of 68,273 infants with up to six measurements in the first year of life. We found dynamic patterns of random effects including time-varying heritability and genetic correlations, as well as several genetic variants showing time-dependent effects, highlighting the applicability of FEMA-Long to enable novel discoveries. FEMA-Long is available at: https://github.com/cmig-research-group/cmig_tools. Author summaryMost large-scale datasets have complexities such as repeated measures, related individuals, or other dependencies across samples, preventing the use of standard regression approaches for analysis. In such circumstances, linear mixed-effects modeling is often employed. However, for high-dimensional datasets, fitting these models is quite challenging. Further, most standard uses of linear mixed-effects modeling focus on simpler covariance models, which may not hold. Here, we introduce FEMA-Long, a novel computationally efficient analytical framework for fitting linear mixed-effects models with time-varying random effects, as well as allowing the effect of the covariates to change smoothly over time by using splines. This is particularly relevant when, for example, studying the effect of genetic variants on phenotypes, where the effects could be non-linear over time. The FEMA-Long framework allows time-varying heritability as well as discovery of genetic variants that show time-dependent effects. By performing a genome-wide association study on data from the Norwegian Mother, Father and Child Cohort Study (MoBa) using FEMA-Long, we show the discovery of genetic variants with time-dependent effects on infant length, weight, and BMI during the first year of life. Our results highlight the potential of using FEMA-Long to make novel discoveries that can lead to biological insights on the genetics of complex traits as well as improve the potential of using genetics for personalized prediction.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Parekh, P., Parker, N., Pecheva, D., Frei, E., Vaudel, M., Smith, D. M., Rigby, A., Jahołkowski, P., Sonderby, I. E., Birkenaes, V., Bakken, N. R., Fan, C. C., Makowski, C., Kopal, J., Loughnan, R. J., Hagler, D. J., van der Meer, D., Johansson, S., Njolstad, P. R., Jernigan, T. L., Thompson, W. K., Frei, O., Shadrin, A. A., Nichols, T. E., Andreassen, O. A., Dale, A. M.. 2025-05-15. FEMA-Long: Modeling unstructured covariances for discovery of time-dependent effects in large-scale longitudinal datasets. https://doi.org/10.1101/2025.05.09.653146

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

The histone demethylase Kdm5 and the ARGONAUTE proteins Piwi and Aubergine regulate female abdominal pigmentation in Drosophila melanogaster

Insect pigmentation is an ecologically critical trait influencing many physiological processes. In Drosophila melanogaster, abdominal pigmentation is sexually dimorphic: males have fully pigmented posterior segments, while females exhibit a posterior melanin stripe. Pigmentation relies on the expression of pigmentation genes that encode enzymes involved in pigment synthesis. These genes are tightly regulated during pupal and young adult stages. To expand the gene regulatory network of pigmentation genes, we conducted an RNAi screen using the yellow-Gal4 driver, expressed during the pupal stage in abdominal epidermis. One of the candidates from this screen, Kdm5, encodes a histone demethylase erasing the H3K4me3 histone mark catalyzed by the histone methyl-transferase Trithorax (Trx). We show that Kdm5 down-regulation reduces abdominal pigmentation, mimicking trx down-regulation. Kdm5 activates melanin production through regulation of the pigmentation gene tan. Transcriptomic analyses reveal that Kdm5 and Trx share many targets in pupal abdominal epidermis, including piRNA pathway components such as piwi and aubergine. These piRNA components, originally associated with transposon silencing in the germline, also function in some somatic tissues such as the nervous system, the fat body or the gut. We demonstrate that Piwi and Aubergine participate in female abdominal pigmentation establishment, without evident piRNA production. We also show that Kdm5 and Piwi act not only in pupal abdominal epidermis but also in pupal fat body. This study therefore expands the regulatory network of pigmentation genes. It identifies a new somatic function for Kdm5 and Piwi and reveals a role for pupal fat body in female abdominal pigmentation regulation.

genetics↗

Genetic diversity within and between polyploid sugarcane (Saccharum spp.) families obtained via caryopsis using microsatellite markers and multicategory model

Genetic diversity analyses are essential for sugarcane (Saccharum spp.) breeding programs. Crossbreeding, based on genetic distances between parental plants, is a tool used to increase genetic variability and enhance plant selection; however, quantifying variation in highly polyploid species remains a challenge. The present study aimed to evaluate the diversity within and between 12 families of sugarcane derived from caryopses, analyzing 120 individual seedlings arranged in an augmented block design. Genotyping was performed using primers for 16 microsatellite loci, five simple sequence repeat (SSR) loci, and 11 expressed sequence tag-SSR (EST-SSR) loci. To accurately account for polyploidy, similarity calculations were performed using Bruvos distances among individuals and RST distances among the families. Analysis of molecular variance (AMOVA) indicated that most of the genetic variability was within families (72%), with only 28% found between them. This high level of intra-family variation demonstrates that a significant reservoir of genetic diversity remains available within the crosses. The highest genetic similarity was observed between the families RB986952 x RB986960 and RB036122 x RB03611, whereas the lowest genetic similarity was observed between the families RB97319 x RB966928 and RB106802 x RB855036. Although the evaluated families shared high genetic similarity, the pronounced genetic variation within them demonstrates a robust recombination potential, indicating that the genetic basis of sugarcane can be better explored using the high variability that already exists in the selection of desirable morpho-agronomic characteristics within the families. Furthermore, this study highlights the importance of using appropriate distances for diversity studies with codominant markers, such as microsatellites, in polyploid species.

genetics↗

Optimizing DNA extraction from environmentally degraded bone samples for molecular identification of cetacean species

Molecular identification of cetacean bone remains can be limited by DNA degradation and the presence of PCR inhibitors. Here, we present an optimized DNA extraction protocol based on a total demineralization method for environmentally exposed cetacean bones. The protocol uses 100 mg of bone powder, 24 h digestion with EDTA, N-lauroylsarcosine, and proteinase K, followed by a modified silica-column purification. Nine environmentally degraded bone samples representing eight individuals were processed. DNA concentrations ranged from 7.3 to 57.1 ng/uL (mean SD = 25.91- 13.91 ng/uL). The mitochondrial cytochrome b gene was successfully amplified from all samples using conventional PCR, and five samples (55.6%) yielded sequences suitable for downstream analysis. BLASTn identified Balaenoptera physalus as the closest database match for all recovered sequences, and phylogenetic analysis further supported their association with B. physalus reference sequences. These results demonstrate that the proposed protocol provides a practical approach for recovering amplifiable and molecularly informative mitochondrial DNA from environmentally degraded cetacean bone material, facilitating molecular identification from challenging skeletal remains.

genetics↗