bioRxiv ScienceSearch

Biology subjects

Huang, X.

Publications and source records attributed to Huang, X..

At least 19 recordsLinked to original sources

Molecular profiles and mutation burden analysis in Chinese patients with gastric carcinoma

The goal of this work was to investigate the molecular profiles and mutation burden in Chinese patients with gastric carcinoma (GC). In total, we performed whole exome sequencing (WES) on 74 GC patients with tumor and adjacent normal formalin-fixed, paraffin-embedded (FFPE) tissue samples. The mutation spectrum of these samples showed a high concordance with TCGA and other studies on GC. We found the alterations of 17 DNA repair genes (including BRCA2, POLE and MSH3, etc.) were strongly correlated with the tumor mutation burden (TMB) and tumor neoantigen burden (TNB) of GC patients. Patients with mutations of these genes tend to have high TMB (median of TMB = 12.77, p=2.3e-6) and TNB (median of TNB = 5.97, p= 2.8e-3). In addition, younger GC patients (age < 60) have lower TMB (p = 0.0021) and TNB (p = 0.034) than older patients (age >= 60). Furthermore, we found a list of 18 genes and two genomic regions (1p36.21 and Xq26.3) were associated with peritoneal metastasis (PM) of GC, and patients with amplification of 1p36.21 and Xq26.3 have a worse prognosis (p=0.002, 0.01, respectively). Our analysis provides GC patients with potential markers for single and combination therapies.

cancer biology

NetGO: Improving Large-scale Protein Function Prediction with Massive Network Information

Automated function prediction (AFP) of proteins is of great significance in biology. In essence, AFP is a large-scale multi-label classification over pairs of proteins and GO terms. Existing AFP approaches, however, have their limitations on both sides of proteins and GO terms. Using various sequence information and the robust learning to rank (LTR) framework, we have developed GOLabeler, a state-of-the-art approach of CAFA3, which overcomes the limitation of the GO term side, such as imbalanced GO terms. Unfortunately, for the protein side issue, available abundant protein information, except for sequences, have not been effectively used for large-scale AFP in CAFA. We propose NetGO that is able to improve large-scale AFP with massive network information. The novelties of NetGO have threefold in using network information: 1) the powerful LTR framework of NetGO efficiently and effectively integrates both sequence and network information, which can easily make large-scale AFP; 2) NetGO can use whole and massive network information of all species (>2000) in STRING (other than only high confidence links and/or some specific species); and 3) NetGO can still use network information to annotate a protein by homology transfer even if it is not covered in STRING. Under numerous experimental settings, we examined the performance of NetGO, such as general performance comparison, species-specific prediction, and prediction on difficult proteins, by using training and test data separated by time-delayed settings of CAFA. Experimental results have clearly demonstrated that NetGO outperforms GOLabeler, DeepGO, and other compared baseline methods significantly. In addition, several interesting findings from our experiments on NetGO would be useful for future AFP research.

bioinformatics

High content analysis methods enable high throughput nematode discovery screening for viability and movement behavior in a multiplex sample in response to natural product treatment.

Monitoring nematode parasite movement and mortality in response to various treatment samples usually involves tedious manual microscopic analysis. High Content Analysis instrumentation enables rapid and high throughput collecting of large numbers of treatment data on huge numbers of individual worms. These large sample sizes and increased sample diversity result in robust, reliable results with increased statistical significance. These methods would be applicable to relevant human, crop, or animal worm parasites.

systems biology

5-Hydroxymethylcytosines from Circulating Cell-free DNA as Diagnostic and Prognostic Markers for Hepatocellular Carcinoma

The lack of highly sensitive and specific diagnostic biomarkers is a major contributor to the poor outcomes of patients with hepatocellular carcinoma (HCC), the second-most common cause of cancer deaths worldwide. We sought to develop a clinically convenient and minimally-invasive approach that can be deployed at scale for the sensitive, specific, and highly reliable diagnosis of HCC, and to evaluate the potential prognostic value of this approach. The study cohort comprised of 2,728 subjects, including HCC patients (n = 1,208), controls (n = 965) (572 healthy individuals and 393 patients with benign lesions), as well as patients with chronic hepatitis B infection (CHB) (n =291), liver cirrhosis (LC) (n = 110), and cholangiocarcinoma (CCC) (n = 154), was recruited from three major liver cancer hospitals in Shanghai, China, from July 2016 to November 2017. Circulating cell-free DNA (cfDNA) were collected from plasma samples from these individuals before surgery or any radical treatment. Applying our 5hmC-Seal technique, the summarized 5-hydroxymethylcytosine (5hmC) profiles in cfDNA were obtained. Molecular annotation analysis suggested that the profiled 5hmC loci in cfDNA were enriched with liver tissue-derived regulatory markers (e.g., H3K4me1). We showed that a weighted diagnostic score (wd-score) based on 117 genes detected using the summarized 5hmC profiles in cfDNA accurately distinguished HCC patients from controls (AUC = 95.1%; 95% CI, 93.6-96.5%) in the validation set, markedly outperformed -fetoprotein (AFP) with superior sensitivity. The wd-scores, which not only detected early BCLC stages (e.g., Stage 0: AUC = 96.2%; 95% CI,94.1-98.4%) and small tumors (e.g., < 2 cm: AUC = 95.7%; 95% CI: 93.6-97.7%), also showed high capacity for distinguishing HCC from non-cancer patients with CHB/LC (AUC = 80.2%; 95% CI, 75.8-84.6%). Moreover, the prognostic value of 5hmC markers in cfDNA was evaluated for HCC recurrence, showing that a weighted prognostic score (wp-score) based on 16 marker genes predicted the recurrence risk (HR = 6.67; 95% CI, 2.81-15.82, p < 0.0001) in 555 patients who have been followed up after surgery. In conclusion, we have developed and validated a robust 5hmC-based diagnostic model that can be applied routinely with clinically feasible amount of cfDNA (e.g., from ~2-5 mL of plasma). Applying this new approach in the clinic could significantly improve the clinical outcomes of HCC patients, for example by early detection of those patients with surgically resectable tumors or as a convenient disease surveillance tool for recurrence.

cancer biology

Genomic sequence capture of haemosporidian parasites: Methods and prospects for enhanced study of host-parasite evolution

Avian malaria and related haemosporidians (Plasmodium, [Para]Haemoproteus, and Leucocytoozoon) represent an exciting multi-host, multi-parasite system in ecology and evolution. Global research in this field accelerated after 1) the publication in 2000 of PCR protocols to sequence a haemosporidian mitochondrial (mtDNA) barcode, and 2) the development in 2009 of an open-access database to document the geographic and host ranges of parasite mtDNA haplotypes. Isolating haemosporidian nuclear DNA from bird hosts, however, has been technically challenging, slowing the transition to genomic-scale sequencing techniques. We extend a recently-developed sequence capture method to obtain hundreds of haemosporidian nuclear loci from wild bird samples, which typically have low levels of infection, or parasitemia. We tested 51 infected birds from Peru and New Mexico and evaluated locus recovery in light of variation in parasitemia, divergence from reference sequences, and pooling strategies. Our method was successful for samples with parasitemia as low as [~]0.03% (3 of 10,000 blood cells infected) and mtDNA divergence as high as 15.9% (one Leucocytozoon sample), and using the most cost-effective pooling strategy tested. Phylogenetic relationships estimated with >300 nuclear loci were well resolved, providing substantial improvement over the mtDNA barcode. We provide protocols for sample preparation and sequence capture including custom probe kit sequences, and describe our bioinformatics pipeline using aTRAM 2.0, PHYLUCE, and custom Perl and Python scripts. This approach can be applied to the tens of thousands of avian samples that have already been screened for haemosporidians, and greatly improve our understanding of parasite speciation, biogeography, and evolutionary dynamics.

genomics

Integrated Analysis Revealed Hub Genes in Breast Cancer

The aim of this study was to identify the hub genes in breast cancer and provide further insight into the tumorigenesis and development of breast cancer. To explore the hub genes in breast cancer, we performed an integrated bioinformatics analysis. Two gene expression profiles were downloaded from the GEO database. The differentially expressed genes (DEGs) were identified by using the \"limma\" package. Then, we performed Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis to explore the functional annotation and potential pathways of the DEGs. Next, protein-protein interaction (PPI) network analysis and weighted gene coexpression network analysis (WGCNA) were conducted to screen for hub genes. To confirm the reliability of the identified hub genes, we obtained TCGA-BRCA data by using WGCNA to screen for genes that were strongly related to breast cancer. By combining the results from the GEO and TCGA datasets, we finally identified 15 real hub genes in breast cancer. Finally, we performed an overall survival analysis to explore the connection between the expression of hub genes and the overall survival time of breast cancer patients. We found that for all hub genes, higher expression was associated with significantly shorter overall survival times among breast cancer patients.

bioinformatics

Hidden Markov Models Lead to Higher Resolution Maps of Mutation Signature Activity in Cancer

Knowing the activity of the mutational processes shaping a cancer genome may provide insight into tumorigenesis and personalized therapy. It is thus important to uncover the characteristic signatures of active mutational processes in patients from their patterns of single base substitutions. However, mutational processes do not act uniformly on the genome and are biased by factors such as the genomes chromatin structure or replication origins. These factors may lead to statistical dependencies among neighboring mutations, calling for modeling approaches that can account for such dependencies to better estimate mutational process activities.\n\nHere we develop the first sequence-dependent models for mutation signatures. We apply these models to characterize genomic and other factors that influence the activity of previously validated mutation signatures in breast cancer. We find that our tool, SO_SCPLOWIGC_SCPLOWMO_SCPLOWAC_SCPLOW, can accurately assign genomic mutations to mutation signatures, yielding assignments that are of higher likelihood than those obtained with models that assume independence between signatures and align better with current biological knowledge. Our analysis resolves a controversy related to the dependency of APOBEC signatures on replication time and links Signatures 18 and 30 to oxidative damage.\n\nModeling the sequential dependencies of mutation signatures leads to improved estimates of mutation signature activity both at the tumor-level and within specific genomic regions, yielding higher resolution maps of mutation signature activity in cancer.

bioinformatics

Characterization of the Rosa roxbunghii Tratt transcriptome and analysis of MYB genes

Rosa roxbunghii Tratt belongs to the Rosaceae family, and the fruit is flavorful, economic, and highly nutritious, providing health benefits. MYB proteins play key roles in R. roxbunghii fruit development and quality. However, the available genomic and transcriptomic information are extremely deficient. Here, a normalized cDNA library was constructed using five tissues, stem, leaf, flower, young fruit, and mature fruit, with three repetitions, and sequenced using the Illumina HiSeq 2500 platform. De novo assembly was performed, and 470.66 million clean reads were obtained. In total, 63,727 unigenes, with an average GC content of 42.08%, were determined and 59,358 were annotated. In addition, 9,354 unigenes were assigned the Gene Ontology category, and 20,202 unigenes were assigned to 25 Eukaryotic Ortholog Groups. Additionally, 19,507 unigenes were classified into 140 pathways of the Kyoto Encyclopedia of Genes and Genomes database. Using the transcriptome, 18 candidate MYB genes that were significantly expressed in mature fruit, compared with other tissues, were obtained. Among them, 10 R2R3 MYB and 1 R1 MYB were identified. The expression levels of 12 MYB genes randomly selected for qRT-PCR analysis were consistent with the RNA-seq results. A total of 37,545 microsatellites were detected, with an average EST--SSR frequency of 0.59 (37,545/63,727). This transcriptome data will be valuable for identifying genes of interest and studying their expression and evolution.

bioinformatics

Identification, Genotyping, and Pathogenicity of Trichosporon spp. Isolated from Giant Pandas

Trichosporon is the dominant genus of epidermal fungi in giant pandas and causes local and deep infections. To provide the information needed for the diagnosis and treatment of trichosporosis in giant pandas, the sequence of ITS, D1/D2, and IGS1 loci in 29 isolates of Trichosporon spp. which isolated from the body surface of giant pandas were combination to investigate interspecies identification and genotype. Morphological development was examined via slide culture. Additionally, mice were infected by skin inunction, intraperitoneal injection, and subcutaneous injection for evaluation of pathogenicity. The twenty-nine isolates of Trichosporon spp. were identified as belonging to 11 species, and Trichosporon jirovecii and T. asteroides were the commonest species. Four strains of T. laibachii and one strain of T. moniliiforme were found to be of novel genotypes, and T. jirovecii was identified to be genotype 1. T. asteroides had the same genotype which involved in disseminated trichosporosis. The morphological development processes of the Trichosporon spp. were clearly different, especially in the processes of single-spore development. Pathogenicity studies showed that 7 species damaged the liver and skin in mice, and their pathogenicity was stronger than other 4 species. T. asteroides had the strongest pathogenicity and might provoke invasive infection. The pathological characteristics of liver and skin infections caused by different Trichosporon spp. were similar. So it is necessary to identify the species of Trichosporon on the surface of giant panda. Combination of ITS, D1/D2, and IGS1 loci analysis, and morphological development process can effectively identify the genotype of Trichosporon spp.

microbiology

Glutathione-S-transferase from the arsenic hyperaccumulator fern Pteris vittata can confer increased arsenate resistance in Escherichia coli

Although arsenic is generally a toxic compound, there are a number of ferns in the genus Pteris that can tolerate large concentrations of this metalloid. In order to probe the mechanisms of arsenic hyperaccumulation, we expressed a Pteris vittata cDNA library in an Escherichia coli {Delta}arsC (arsenate reductase) mutant. We obtained three independent clones that conferred increased arsenate resistance on this host. DNA sequence analysis indicated that these clones specify proteins that have a high sequence similarity to the phi class of glutathione-S-transferases (GSTs) of higher plants. Detoxification of arsenate by the P. vittata GSTs in E. coli was abrogated by a gshA mutation, which blocks the synthesis of glutathione, and by a gor mutation, which inactivates glutathione reductase. Direct measurements of the speciation of arsenic in culture media of the E. coli strains expressing the P. vittata GSTs indicated that these proteins facilitate the reduction of arsenate. Our observations suggest that the detoxification of arsenate by the P. vittata GSTs involves reduction of As(V) to As(III) by glutathione or a related sulfhydro compound.\n\nFundingThe authors acknowledge support from the Indiana 21st Century Research and technology Fund (912010479) to DES and LNC, from the U.S. Department of Energy (grant no. DE-FG02-03ER63622) to DES, and from BBSRC-DFID (grant no. BBF0041841GJN) to AAM. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. There are no financial, personal, or professional interests that could be construed to have influenced the paper.

plant biology

Overdosage of balanced protein complexes reduces proliferation rate in aneuploid cells

Cells with complex aneuploidies, such as tumor cells, display a wide range of phenotypic abnormalities. However, molecular basis for this has been mainly studied in trisomic (2n+1) and disomic (n+1) cells. To determine how karyotype affects proliferation rate in cells with complex aneuploidies we generated forty 2n+x yeast strains in which each diploid cell has an extra 5 to 12 chromosomes and found that these strains exhibited abnormal cell-cycle progression. Proliferation rate was negatively correlated with the number of protein complexes in which all subunits were at the 3-copy level, but not with the number of imbalanced complexes made up of a mixture of 2-copy and 3-copy genes. Proteomics revealed that most 3-copy members of imbalanced complexes were expressed at only 2n protein levels whereas members of complexes in which all subunits are stoichiometrically balanced at 3 copies per cell had 3n protein levels. We identified individual protein complexes for which overdosage reduces proliferation rate, and found that deleting one copy of each member partially restored proliferation rate in cells with complex aneuploidies. Lastly, we validated this finding using orthogonal datasets from both yeast and from human cancers. Taken together, our study provides a novel explanation how aneuploidy affects phenotype.

systems biology

Determinants of early afterdepolarization properties in ventricular myocyte models

Early afterdepolarizations (EADs) are spontaneous depolarizations during the repolarization phase of an action potential in cardiac myocytes. It is widely known that EADs are promoted by increasing inward currents and/or decreasing outward currents, a condition called reduced repolarization reserve. Recent studies based on bifurcation theories show that EADs are caused by a dual Hopf-homoclinic bifurcation, bringing in further mechanistic insights into the genesis and dynamics of EADs. In this study, we investigated the EAD properties, such as the EAD amplitude, the inter-EAD interval, and the latency of the first EAD, and their major determinants. We first made predictions based on the bifurcation theory and then validated them in physiologically more detailed action potential models. These properties were investigated by varying one parameter at a time or using parameter sets randomly drawn from assigned intervals. The theoretical and simulation results were compared with experimental data from the literature. Our major findings are that the EAD amplitude and takeoff potential exhibit a negative linear correlation; the inter-EAD interval is insensitive to the maximum ionic current conductance but mainly determined by the kinetics of ICa,L and the dual Hopf-homoclinic bifurcation; and both inter-EAD interval and latency vary largely from model to model. Most of the model results generally agree with experimental observations in isolated ventricular myocytes. However, a major discrepancy between modeling results and experimental observations is that the inter-EAD intervals observed in experiments are mainly between 200 and 500 ms, irrespective of species, while those of the mathematical models exhibit a much wider range with some models exhibiting inter-EAD intervals less than 100 ms. Our simulations show that the cause of this discrepancy is likely due to the difference in ICa,L recovery properties in different mathematical models, which needs to be addressed in future action potential model development.\n\nAuthor summaryEarly afterdepolarizations (EADs) are abnormal depolarizations during the plateau phase of action potential in cardiac myocytes, arising from a dual Hopf-homoclinic bifurcation. The same bifurcations are also responsible for certain types of bursting behaviors in other cell types, such as beta cells and neuronal cells. EADs are known to play important role in the genesis of lethal arrhythmias and have been widely studied in both experiments and computer models. However, a detailed comparison between the properties of EADs observed in experiments and those from mathematical models have not been carried out. In this study, we performed theoretical analyses and computer simulations of different ventricular action potential models as well as different species to investigate the properties of EADs and compared these properties to those observed in experiments. While the EAD properties in the action potential models capture many of the EAD properties seen in experiments, the inter-EAD intervals in the computer models differ a lot from model to model, and some of them show very large discrepancy with those observed in experiments. This discrepancy needs to be addressed in future cardiac action potential model development.

biophysics

Base pair editing of goat embryos: nonsense codon introgression into FGF5 to improve cashmere yield

The ability to alter single bases without DNA double strand breaks provides a potential solution for multiplex editing of livestock genomes for quantitative traits. Here, we report using a single base editing system, Base Editor 3 (BE3), to induce nonsense codons (C-to-T transitions) at four target sites in caprine FGF5. All five progenies produced from microinjected single-cell embryos had alleles with a targeted nonsense mutation and yielded expected phenotypes. The effectiveness of BE3 to make single base changes varied considerably based on sgRNA design. Also, the rate of mosaicism differed between animals, target sites, and tissue type. PCR amplicon and whole genome resequencing analyses for off-target changes caused by BE3 were low at a genome-wide scale. This study provides first evidence of base editing in livestock, thus presenting a potentially better method to introgress complex human disease alleles into large animal models and provide genetic improvement of complex health and production traits in a single generation.

genetics

Identification of genes affecting saturated fat acid content in Elaeis guineensis by genome-wide association analysis

Oil palm is the highest yielding oil crop per unit area worldwide. Unfortunately, palm oil is often considered unhealthy. In particular, palmic acid (C16:0) is a major component of palm oil. In this study a total of 1 261 501 SNP markers were produced in a diversity panel of 200 oil palm individuals. Oil content in this population varied from 29.8% to 70.3%, palmic acid varied from 31.3% to 48.8%, and oleic acid varied from 31.3% to 50.1%. We identified 274 SNP markers significantly associated with fatty acid compositions; 44 candidate genes in the flanking regions of these SNPs were involved in fatty acid biosynthesis and metabolism. Among them, two acyl-ACP thioesterase B genes had differential expression patterns between the mesocarp and kernel, tissues which show different oil profiles in oil palm (high palmic acid and high lauric acid respectively). Overexpression of both genes caused a significant increase in palmic acid content, while overexpression of the EgFatB2 gene also caused an accumulation of lauric acid and myristic acid. Our research provides genome-wide SNPs, a set of markers significantly associated with fatty acid content, and validated candidate genes for future targeted breeding of lower saturated fat content in palm oil.

plant biology

First report and multilocus genotyping of Enterocytozoon bieneusi from Tibetan pigs in southwestern China

Enterocytozoon bieneusi is a common intestinal pathogen and a major cause of diarrhea and enteric diseases in a variety of animals. While the E. bieneusi genotype has become better-known, there are few reports on its prevalence in the Tibetan pig. This study investigated the prevalence, genetic diversity, and zoonotic potential of E. bieneusi in the Tibetan pig in southwestern China. Tibetan pig feces (266 samples) were collected from three sites in the southwest of China. Feces were subjected to PCR amplification of the internal transcribed spacer (ITS) region. E. bieneusi was detected in 83 (31.2%) of Tibetan pigs from the three different sites, with 25.4% in Kangding, 56% in Yaan and 26.7% in Qionglai. Age group demonstrated the prevalence of E. bieneusi range from 24.4%(aged 0 to 1 years) to 44.4%(aged 1 to 2 years). Four genotypes of E. bieneusi were identified: two known genotypes EbpC (n=58), Henan-IV (n=24) and two novel genotypes, SCT01 and SCT02 (one of each). Phylogenetic analysis showed these four genotypes clustered to group 1 with zoonotic potential. Multilocus sequence typing (MLST) analysis three microsatellites (MS1, MS3, MS7) and one minisatellite (MS4) revealed 47, 48, 23 and 47 positive specimens were successfully sequenced, and identified ten, ten, five and five genotypes at four loci, respectively. This study indicates the potential danger of E. bieneusi to Tibetan pigs in southwestern China, and offers basic data for preventing and controlling infections.

genetics

Automated Tracking of Biopolymer Growth and Network Deformation with TSOAX

Studies of how individual semi-flexible biopolymers and their network assemblies change over time reveal dynamical and mechanical properties important to the understanding of their function in tissues and living cells. Automatic tracking of biopolymer networks from fluorescence microscopy time-lapse sequences facilitates such quantitative studies. We present an open source software tool that combines a global and local correspondence algorithm to track biopolymer networks in 2D and 3D, using stretching open active contours. We demonstrate its application in fully automated tracking of elongating and intersecting actin filaments, detection of loop formation and constriction of tilted contractile rings in live cells, and tracking of network deformation under shear deformation.

cell biology

RNAs as proximity labeling media for identifying nuclear speckle positions relative to the genome

Nuclear speckles are interchromatin structures enriched in RNA splicing factors. Determining their relative positions with respect to the folded nuclear genome could provide critical information on co-and post-transcriptional regulation of gene expression. However, it remains challenging to identify which parts of the nuclear genome are in proximity to nuclear speckles, due to physical separation between nuclear speckle cores and chromatin. We hypothesized that noncoding RNAs including small nuclear RNAs, 7SK and Malat1, which accumulate at the periphery of nuclear speckles (nsaRNA, nuclear speckle associated RNA), may extend to sufficient proximity to the nuclear genome. Leveraging a transcriptome-genome interaction assay (MARGI), we identified nsaRNA-interacting genomic sequences, which exhibited clustering patterns (nsaPeaks) in the genome, suggesting existence of relatively stable interaction sites for nsaRNAs in nuclear genome. Posttranscriptional pre-mRNAs, which are known to be clustered to nuclear speckles, exhibited proximity to nsaPeaks but rarely to other genomic regions. Furthermore, CDK9 proteins that localize to the vicinity of nuclear speckles produced ChIP-seq peaks that overlapped with nsaPeaks. Our combined DNA FISH and immunofluorescence analysis in 182 single cells revealed a 3-fold increase in odds for nuclear speckles to localize near an nsaPeak than its neighboring genomic sequence. These data suggest a model that nsaRNAs locate in sufficient proximity to nuclear genome and leave identifiable genomic footprints, thus revealing the parts of genome proximal to nuclear speckles.

bioinformatics

Accurate functional classification of thousands of BRCA1 variants with saturation genome editing

Variants of uncertain significance (VUS) fundamentally limit the utility of genetic information in a clinical setting. The challenge of VUS is epitomized by BRCA1, a tumor suppressor gene integral to DNA repair and genomic stability. Germline BRCA1 loss-of-function (LOF) variants predispose women to early-onset breast and ovarian cancers. Although BRCA1 has been sequenced in millions of women, the risk associated with most newly observed variants cannot be definitively assigned. Data sharing attenuates this problem but it is unlikely to solve it, as most newly observed variants are exceedingly rare. In lieu of genetic evidence, experimental approaches can be used to functionally characterize VUS. However, to date, functional studies of BRCA1 VUS have been conducted in a post hoc, piecemeal fashion. Here we employ saturation genome editing to assay 96.5% of all possible single nucleotide variants (SNVs) in 13 exons that encode functionally critical domains of BRCA1. Our assay measures cellular fitness in a haploid human cell line whose survival is dependent on intact BRCA1 function. The resulting function scores for nearly 4,000 SNVs are bimodally distributed and almost perfectly concordant with established assessments of pathogenicity. Sequence-function maps enhanced by parallel measurements of variant effects on mRNA levels reveal mechanisms by which loss-of-function SNVs arise. Hundreds of missense SNVs critical for protein function are identified, as well as dozens of exonic and intronic SNVs that compromise BRCA1 function by disrupting splicing or transcript stability. We predict that these function scores will be directly useful for the clinical interpretation of cancer risk based on BRCA1 sequencing. Furthermore, we propose that this paradigm can be extended to overcome the challenge of VUS in other genes in which genetic variation is clinically actionable.

genomics