bioRxiv ScienceSearch

bioRxiv · 10.1101/741694

Gender bias in research teams and the underrepresentation of women in science

Abstract

Why are females still underrepresented in science? The social factors that affect career choices and trajectories are thought to be important but are poorly understood. We analyzed author gender in a sample of >61,000 scientific articles in the biological sciences to evaluate the factors that shape the formation of research teams. We find that authorship teams are more gender-assorted than expected by chance, with excess homotypic assortment accounting for up to 7% of published articles. One possible mechanism that could explain gender assortment and broader patterns of female representation is that women may focus on different research topics than men (i.e., the \"topic preference\" hypothesis). An alternative hypothesis is that researchers may consciously or unconsciously prefer to work within same-gender teams (the \"gender homophily\" hypothesis). Using network analysis, we find no evidence to support the topic preference hypothesis, because the topics of female-authored articles are no more similar to each other than expected within the broader research landscape. Instead, consistent with a model of moderate gender homophily, we find that the prevalence of matched-gender teams increases as a discipline moves towards gender parity. This can occur because latent preferences are more easily fulfilled in a gender-diverse environment. Finally, we show that female authors pay a substantial citation cost to work in gender-matched teams. Notably, the prevalence of homotypic assortment is predicted to increase in the future if more disciplines shift towards gender parity. These data indicate that social preferences can have important downstream consequences for the retention of women and other underrepresented groups in science.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Dakin, R., Ryder, T. B.. 2019-08-24. Gender bias in research teams and the underrepresentation of women in science. https://doi.org/10.1101/741694

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education

PlotTwist - a web app for plotting and annotating time-series data

The results from time-dependent experiments are often used to generate plots that visualize how the data evolves over time. To simplify state-of-the-art data visualization and annotation of data from such experiments, an open source tool was created with R/shiny that does not require coding skills to operate. The freely available web app accepts wide (spreadsheet) and tidy data and offers a range of options to normalize the data. The data from individual objects can be shown in three different ways: (i) lines with unique colors, (ii) small multiples and (iii) heatmap-style display. Next to this, the mean can be displayed with a 95% confidence interval for the visual comparison of different conditions. Several color blind friendly palettes are available to label the data and/or statistics. The plots can be annotated with graphical features and/or text to indicate any perturbations that were applied during the time-lapse experiments. All user-defined settings can be stored for reproducibility of the data visualization. The app is dubbed PlotTwist and is available online: https://huygens.science.uva.nl/PlotTwist\n\n\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=166 SRC=\"FIGDIR/small/745612v2_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (34K):\norg.highwire.dtl.DTLVardef@1bad5e6org.highwire.dtl.DTLVardef@1312023org.highwire.dtl.DTLVardef@350cb5org.highwire.dtl.DTLVardef@d556c6_HPS_FORMAT_FIGEXP M_FIG C_FIG

scientific communication and education

The Five-Primer Challenge: An inquiry-based laboratory module for synthetic biology

New technologies in DNA synthesis and assembly give genetic engineers complete freedom in genetic design, where virtually any plasmid DNA sequence can be created efficiently and economically. Learning how to design, construct, and test new DNA sequences is a critical skill for researchers in molecular biology and biotechnology. Here we present a student-centered, inquiry-based module in which students learn how to control bacterial gene expression by appplying various DNA assembly techniques. The central activity in this learning module is termed the Five-Primer Challenge. Each student is allowed to order up to five 60-mer oligonucleotide primers to then modify a GFP expression plasmid with the goal of increasing GFP expression as much as possible. This module was developed and implemented at the 2016 Cold Spring Harbor Laboratory Synthetic Biology Course, and was effective at engaging students in critical thinking and in promoting student learning.

scientific communication and education