bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

Multiple reference genome sequences of hot pepper reveal the massive evolution of plant disease resistance genes by retroduplication

Transposable elements (TEs) provide major evolutionary forces leading to new genome structure and species diversification. However, the role of TEs in the expansion of disease resistance gene families has been unexplored in plants. Here, we report high-quality de novo genomes for two peppers (Capsicum baccatum and C. chinense) and an improved reference genome (C. annuum). Dynamic genome rearrangements involving translocations among chromosome 3, 5 and 9 were detected in comparison between C. baccatum and the two other peppers. The amplification of athila LTR-retrotransposons, members of the gypsy superfamily, led to genome expansion in C. baccatum. In-depth genome-wide comparison of genes and repeats unveiled that the copy numbers of NLRs were greatly increased by LTR-retrotransposon-mediated retroduplication. Moreover, retroduplicated NLRs exhibited great abundance across the angiosperms, with most cases lineage-specific and thus recent events. Our study revealed that retroduplication has played key roles in the emergence of new disease-resistance genes in plants.

plant biology

Estimation Of Genomic Prediction Accuracy From Reference Populations With Varying Degrees Of Relationship

Genomic prediction is emerging in a wide range of fields including animal and plant breeding, risk prediction in human precision medicine and forensic. It is desirable to establish a theoretical framework for genomic prediction accuracy when the reference data consists of information sources with varying degrees of relationship to the target individuals. A reference set can contain both close and distant relatives as well as unrelated individuals from the wider population in the genomic prediction. The various sources of information were modeled as different populations with different effective population sizes (Ne). Both the effective number of chromosome segments (Me) and Ne are considered to be a function of the data used for prediction. We validate our theory with analyses of simulated as well as real data, and illustrate that the variation in genomic relationships with the target is a predictor of the information content of the reference set. With a similar amount of data available for each source, we show that close relatives can have a substantially larger effect on genomic prediction accuracy than lesser related individuals. We also illustrate that when prediction relies on closer relatives, there is less improvement in prediction accuracy with an increase in training data or marker panel density. We release software that can estimate the expected prediction accuracy and power when combining different reference sources with various degrees of relationship to the target, which is useful when planning genomic prediction (before or after collecting data) in animal, plant and human genetics.

genetics

Assembly Of Whole-Chromosome Pseudomolecules For Polyploid Plant Genomes Using Outcrossed Mapping Populations

The assembly of whole-chromosome pseudomolecules for plant genomes remains challenging due to polyploidy and high repeat content. We developed an approach for constructing complete pseudomolecules for polyploid species using genotyping-by-sequencing data from outcrossing mapping populations coupled with high coverage whole genome sequence data of a reference genome. Our approach combines de novo assembly with linkage mapping to arrange scaffolds into pseudomolecules. We show that the method is able to reconstruct simulated chromosomes for both diploid and tetraploid genomes. Comparisons to three existing genetic mapping tools show that our method outperforms the other methods in accuracy on both grouping and ordering, and is robust to the presence of substantial amounts of missing data and genotyping errors. We applied our method to three real datasets including a diploid Ipomoea trifida and two tetraploid potato mapping populations. The linkage maps show significant concordance with the reference chromosomes. We resolved seven assembly errors for the published Ipomoea trifida genome assembly as well as anchored an unplaced scaffold in the published potato genome.

bioinformatics

Genomics-enabled analysis of the emergent disease cotton bacterial blight

Cotton bacterial blight (CBB), an important disease of (Gossypium hirsutum) in the early 20th century, had been controlled by resistant germplasm for over half a century. Recently, CBB re-emerged as an agronomic problem in the United States. Here, we report analysis of cotton variety planting statistics that indicate a steady increase in the percentage of susceptible cotton varieties grown each year since 2009. Phylogenetic analysis revealed that strains from the current outbreak cluster with race 18 Xanthomonas citri pv. malvacearum (Xcm) strains.\n\nIllumina based draft genomes were generated for thirteen Xcm isolates and analyzed along with 4 previously published Xcm genomes. These genomes encode 24 conserved and nine variable type three effectors. Strains in the race 18 clade contain 3 to 5 more effectors than other Xcm strains. SMRT sequencing of two geographically and temporally diverse strains of Xcm yielded circular chromosomes and accompanying plasmids. These genomes encode eight and thirteen distinct transcription activator-like effector genes. RNA-sequencing revealed 52 genes induced within two cotton cultivars by both tested Xcm strains. This gene list includes a homeologous pair of genes, with homology to the known susceptibility gene, MLO. In contrast, the two strains of Xcm induce different clade III SWEET sugar transporters. Subsequent genome wide analysis revealed patterns in the overall expression of homeologous gene pairs in cotton after inoculation by Xcm. These data reveal host-pathogen specificity at the genetic level and strategies for future development of resistant cultivars.\n\nAuthor SummaryCotton bacterial blight (CBB), caused by Xanthomonas citri pv. malvacearum (Xcm), significantly limited cotton yields in the early 20th century but has been controlled by classical resistance genes for more than 50 years. In 2011, the pathogen re-emerged with a vengeance. In this study, we compare diverse pathogen isolates and cotton varieties to further understand the virulence mechanisms employed by Xcm and to identify promising resistance strategies. We generate fully contiguous genome assemblies for two diverse Xcm strains and identify pathogen proteins used to modulate host transcription and promote susceptibility. RNA-Sequencing of infected cotton reveals novel putative gene targets for the development of durable Xcm resistance. Together, the data presented reveal contributing factors for CBB re-emergence in the U.S. and highlight several promising routes towards the development of durable resistance including classical resistance genes and potential manipulation of susceptibility targets.

plant biology

Improved DOP-PCR (iDOP-PCR): a robust and simple WGA method for efficient amplification of low copy number genomic DNA

Whole-genome amplification (WGA) techniques are used for non-specific amplification of low-copy number DNA, and especially for single-cell genome and transcriptome amplification. There are a number of WGA methods that have been developed over the years. One example is degenerate oligonucleotide-primed PCR (DOP-PCR), which is a very simple, fast and inexpensive WGA technique. Although DOP-PCR has been regarded as one of the pioneering methods for WGA, it only provides low genome coverage and a high allele dropout rate when compared to more modern techniques. Here we describe an improved DOP-PCR (iDOP-PCR). We have modified the classic DOP-PCR by using a new thermostable DNA polymerase (SD polymerase) with a strong strand-displacement activity and by adjustments in primers design. We compared iDOP-PCR, classic DOP-PCR and the well-established PicoPlex technique for whole genome amplification of both high- and low-copy number human genomic DNA. The amplified DNA libraries were evaluated by analysis of short tandem repeat genotypes and NGS data. In summary, iDOP-PCR provided a better quality of the amplified DNA libraries compared to the other WGA methods tested, especially when low amounts of genomic DNA were used as an input material.

molecular biology

Estimation Of Universal And Taxon-Specific Parameters Of Prokaryotic Genome Evolution

Our recent study on mathematical modeling of microbial genome evolution indicated that, on average, genomes of bacteria and archaea evolve in the regime of mutation-selection balance defined by positive selection coefficients associated with gene acquisition that is counter-acted by the intrinsic deletion bias. This analysis was based on the strong assumption that parameters of genome evolution are universal across the diversity of bacteria and archaea, and yielded extremely low values of the selection coefficient. Here we further refine the modeling approach by taking into account evolutionary factors specific for individual groups of microbes using two independent fitting strategies, an ad hoc hard fitting scheme and an hierarchical Bayesian model. The resulting estimate of the mean selection coefficient of s[~]10-10 associated with the gain of one gene implies that, on average, acquisition of a gene is beneficial, and that microbial genomes typically evolve under a weak selection regime that might transition to strong selection in highly abundant organisms with large effective population sizes. The apparent selective pressure towards larger genomes is balanced by the deletion bias, which is estimated to be consistently greater than unity for all analyzed groups of microbes. The estimated values of s are more realistic than the lower values obtained previously, indicating that global and group-specific evolutionary factors synergistically affect microbial genome evolution that seems to be driven primarily by adaptation to existence in diverse niches.

evolutionary biology

Comparative Genomics Sheds Light On Niche Differentiation And The Evolutionary History Of Comammox Nitrospira

The description of comammox Nitrospira spp., performing complete ammonium-to-nitrate oxidation, and their co-occurrence with canonical betaproteobacterial ammonium oxidizing bacteria ({beta}-AOB) in the environment, call into question the metabolic potential of comammox Nitrospira and the evolutionary history of their ammonium oxidation pathway. We report four new comammox Nitrospira genomes, constituting two novel species, and the first comparative genomic analysis on comammox Nitrospira.\n\nComammox Nitrospira has lost the potential to use external nitrite as energy and nitrogen source: compared to strictly nitrite oxidizing Nitrospira; they lack genes for assimilative nitrite reduction and reverse electron transport from nitrite. By contrast, compared to other Nitrospira, their ammonium oxidizer physiology is exemplified by genes for ammonium and urea transporters and copper homeostasis and the lack of cyanate hydratase genes. Two comammox clades are different in their ammonium uptake systems. Contrary to {beta}-AOB, comammox Nitrospira genomes have single copies of the two central ammonium oxidation pathway genes, lack genes involved in nitric oxide reduction, and encode genes that would allow efficient growth at low oxygen concentrations. Hence, comammox Nitrospira seems attuned to oligotrophy and hypoxia compared to {beta}-AOB.\n\n{beta}-AOBs are the clear origin of the ammonium oxidation pathway in comammox Nitrospira: reconciliation analysis indicates two separate early amoA gene transfer events from {beta}-AOB to an ancestor of comammox Nitrospira, followed by clade specific losses. For haoA, one early transfer from {beta}-AOB to comammox Nitrospira is predicted - followed by intra-clade transfers. We postulate that the absence of comammox genes in most Nitrospira genomes is the result of subsequent loss.\n\nSignificanceThe recent discovery of comammox bacteria - members of the Nitrospira genus able to fully oxidize ammonia to nitrate - upset the long-held conviction that nitrification is a two-step process. It also opened key questions on the ecological and evolutionary relations of these bacteria with other nitrifying prokaryotes. Here, we report the first comparative genomic analysis of comammox Nitrospira and related nitrifiers. Ammonium oxidation genes in comammox Nitrospira had a surprisingly complex evolution, originating from ancient transfer from the phylogenetically distantly related ammonia-oxidizing betaproteobacteria, followed by within-lineage transfers and losses. The resulting comammox genomes are uniquely adapted to ammonia oxidation in nutrient-limited and low-oxygen environments and appear to have lost the genetic potential to grow by nitrite oxidation alone.

microbiology

Genome-Enabled Insights Into The Ecophysiology Of The Comammox Bacterium Candidatus Nitrospira nitrosa

The recently discovered comammox bacteria have the potential to completely oxidize ammonia to nitrate. These microorganisms are part of the Nitrospira genus and are present in a variety of environments, including Biological Nutrient Removal (BNR) systems. However, the physiological traits within and between comammox- and nitrite oxidizing bacteria (NOB)-like Nitrospira species have not been analyzed in these ecosystems. In this study, we identified Nitrospira strains dominating the nitrifying community of a sequencing batch reactor (SBR) performing BNR under micro-aerobic conditions. We recovered metagenomes-derived draft genomes from two Nitrospira strains: (1) Nitrospira sp. UW-LDO-01, a comammox-like organism classified as Candidatus Nitrospira nitrosa, and (2) Nitrospira sp. UW-LDO-02, a nitrite oxidizing strain belonging to the Nitrospira defluvii species. A comparative genomic analysis of these strains with other Nitrospira-like genomes identified genomic differences in Ca. Nitrospira nitrosa mainly attributed to each strains niche adaptation. Traits associated with energy metabolism also differentiate comammox from NOB-like genomes. We also identified several transcriptionally regulated adaptive traits, including stress tolerance, biofilm formation and micro-aerobic metabolism, which might explain survival of Nitrospira under multiple environmental conditions. Overall, our analysis expanded our understanding of the genetic functional features of Ca. Nitrospira nitrosa, and identified genomic traits that further illuminate the phylogenetic diversity and metabolic plasticity of the Nitrospira genus.

microbiology

Fine scale mapping of genomic introgressions within the Drosophila yakuba clade

The process of speciation involves populations diverging over time until they are genetically and reproductively isolated. Hybridization between nascent species was long thought to directly oppose speciation. However, the amount of interspecific genetic exchange (introgression) mediated by hybridization remains largely unknown, although recent progress in genome sequencing has made measuring introgression more tractable. A natural place to look for individuals with admixed ancestry (indicative of introgression) is in regions where species co-occur. In west Africa, D. santomea and D. yakuba hybridize on the island of Sao Tome, while D. yakuba and D. teissieri hybridize on the nearby island of Bioko. In this report, we quantify the genomic extent of introgression between the three species of the Drosophila yakuba clade (D. yakuba, D. santomea), D. teissieri). We sequenced the genomes of 86 individuals from all three species. We also developed and applied a new statistical framework, using a hidden Markov approach, to identify introgression. We found that introgression has occurred between both species pairs but most introgressed segments are small (on the order of a few kilobases). After ruling out the retention of ancestral polymorphism as an explanation for these similar regions, we find that the sizes of introgressed haplotypes indicate that genetic exchange is not recent (>1,000 generations ago). We additionally show that in both cases, introgression was rarer on X chromosomes than on autosomes which is consistent with sex chromosomes playing a large role in reproductive isolation. Even though the two species pairs have stable contemporary hybrid zones, providing the opportunity for ongoing gene flow, our results indicate that genetic exchange between these species is currently rare.\n\nAUTHOR SUMMARYEven though hybridization is thought to be pervasive among animal species, the frequency of introgression, the transfer of genetic material between species, remains largely unknown. In this report we quantify the magnitude and genomic distribution of introgression among three species of Drosophila that encompass the two known stable hybrid zones in this genetic model genus. We obtained whole genome sequences for individuals of the three species across their geographic range (including their hybrid zones) and developed a hidden Markov model-based method to identify patterns of genomic introgression between species. We found that nuclear introgression is rare between both species pairs, suggesting hybrids in nature rarely successfully backcross with parental species. Nevertheless, some D. santomea alleles introgressed into D. yakuba have spread from Sao Tome to other islands in the Gulf of Guinea where D. santomea is not found. Our results indicate that in spite of contemporary hybridization between species that produces fertile hybrids, the rates of gene exchange between species are low.

evolutionary biology

Genomic variations in paired normal controls for lung adenocarcinomas

Somatic genomic mutations in lung adenocarcinomas (LUADs) have been extensively dissected, but whether the counterpart normal lung tissues that are exposed to ambient air or tobacco smoke as the tumor tissues do, harbor genomic variations, remains unclear. Here, the genome of normal lung tissues and paired tumors of 11 patients with LUAD were sequenced, the genome sequences of counterpart normal controls (CNCs) and tumor tissues of 513 patients were downloaded from TCGA database and analyzed. In the initial screening, genomic alterations were identified in the \"normal\" lung tissues and verified by Sanger capillary sequencing. In CNCs of TCGA datasets, a mean of 0.2721 exonic variations/Mb and 5.2885 altered genes per sample were uncovered. The C:G[->]T:A transitions, a signature of tobacco carcinogen N-methyl-N-nitro-N-nitrosoguanidine, were the predominant nucleotide changes in CNCs. 16 genes had a variant rate of more than 2%, and CNC variations in MUC5B, ZXDB, PLIN4, CCDC144NL, CNTNAP3B, and CCDC180 were associated with poor prognosis whereas alterations in CHD3 and KRTAP5-5 were associated with favorable clinical outcome of the patients. This study identified the genomic alterations in CNC samples of LUADs, and further highlighted the DNA damage effect of tobacco on lung epithelial cells.

cancer biology

The genomic architecture of a rapid island radiation: mapping chromosomal rearrangements and recombination rate variation in Laupala

Phenotypic evolution and speciation depend on recombination in many ways. Within populations, recombination can promote adaptation by bringing together favorable mutations and decoupling beneficial and deleterious alleles. As populations diverge, cross-over can give rise to maladapted recombinants and impede or reverse diversification. Suppressed recombination due to genomic rearrangements, modifier alleles, and intrinsic chromosomal properties may offer a shield against maladaptive gene flow eroding co-adapted gene complexes. Both theoretical and empirical results support this relationship. However, little is known about this relationship in the context of behavioral isolation, where co-evolving signals and preferences are the major hybridization barrier. Here we examine the genomic architecture of recently diverged, sexually isolated Hawaiian swordtail crickets (Laupala). We assemble a de novo genome and generate three dense linkage maps from interspecies crosses. In line with expectations based on the species recent divergence and successful interbreeding in the lab, the linkage maps are highly collinear and show no evidence for large-scale chromosomal rearrangements. The maps were then used to anchor the assembly to pseudomolecules and estimate recombination rates across the genome. We tested the hypothesis that loci involved in behavioral isolation (song and preference divergence) are in regions of low interspecific recombination. Contrary to our expectations, a genomic region where a male song QTL co-localizes with a female preference QTL was not associated with particularly low recombination rates. This study provides important novel genomic resources for an emerging evolutionary genetics model system and suggests that trait-preference co-evolution is not necessarily facilitated by locally suppressed recombination.

evolutionary biology

Long-term adaptive evolution of genomically recoded Escherichia coli

Efforts are underway to construct several recoded genomes anticipated to exhibit multi-virus resistance, enhanced non-standard amino acid (NSAA) incorporation, and capability for synthetic biocontainment. Though we succeeded in pioneering the first genomically recoded organism (Escherichia coli strain C321.{Delta}A), its fitness is far lower than that of its non-recoded ancestor, particularly in defined media. This fitness deficit severely limits its utility for NSAA-linked applications requiring defined media such as live cell imaging, metabolic engineering, and industrial-scale protein production. Here, we report adaptive evolution of C321.{Delta}A for more than 1,000 generations in independent replicate populations grown in glucose minimal media. Evolved recoded populations significantly exceed the growth rates of both the ancestral C321.{Delta}A and non-recoded strains, permitting use of the recoded chassis in several new contexts. We use next-generation sequencing to identify genes mutated in multiple independent populations, and we reconstruct individual alleles in ancestral strains via multiplex automatable genome engineering (MAGE) to quantify their effects on fitness. Several selective mutations occur only in recoded evolved populations, some of which are associated with altering the translation apparatus in response to recoding, whereas others are not apparently associated with recoding, but instead correct for off-target mutations that occurred during initial genome engineering. This report demonstrates that laboratory evolution can be applied after engineering of recoded genomes to streamline fitness recovery compared to application of additional targeted engineering strategies that may introduce further unintended mutations. In doing so, we provide the most comprehensive insight to date into the physiology of the commonly used C321.{Delta}A strain.\n\nSignificance StatementAfter demonstrating construction of an organism with an altered genetic code, we sought to evolve this organism for many generations to improve its fitness and learn what unique changes natural selection would bestow upon it. Although this organism initially had impaired fitness, we observed that adaptive laboratory evolution resulted in several selective mutations that corrected for insufficient translation termination and for unintended mutations that occurred when originally altering the genetic code. This work further bolsters our understanding of the pliability of the genetic code, it will help guide ongoing and future efforts seeking to recode genomes, and it results in a useful strain for non-standard amino acid incorporation in numerous contexts relevant for research and industry.

evolutionary biology

NmeCas9 is an intrinsically high-fidelity genome editing platform

BackgroundThe development of CRISPR genome editing has transformed biomedical research. Most applications reported thus far rely upon the Cas9 protein from Streptococcus pyogenes SF370 (SpyCas9). With many RNA guides, wild-type SpyCas9 can induce significant levels of unintended mutations at near-cognate sites, necessitating substantial efforts toward the development of strategies to minimize off-target activity. Although the genome-editing potential of thousands of other Cas9 orthologs remains largely untapped, it is not known how many will require similarly extensive engineering to achieve single-site accuracy within large (e.g. mammalian) genomes. In addition to its off-targeting propensity, SpyCas9 is encoded by a relatively large (~4.2 kb) open reading frame, limiting its utility in applications that require size-restricted delivery strategies such as adeno-associated virus vectors. In contrast, some genome-editing-validated Cas9 orthologs (e.g. from Staphylococcus aureus, Campylobacter jejuni, Geobacillus stearothermophilus and Neisseria meningitidis) are considerably smaller and therefore better suited for viral delivery.\n\nResultsHere we show that wild-type NmeCas9, when programmed with guide sequences of natural length (24 nucleotides), exhibits a nearly complete absence of unintended editing in human cells, even when targeting sites that are prone to off-target activity with wildtype SpyCas9. We also validate at least six variant protospacer adjacent motifs (PAMs), in addition to the preferred consensus PAM (5-N4GATT-3), for NmeCas9 genome editing in human cells.\n\nConclusionsOur results show that NmeCas9 is a naturally high-fidelity genome editing enzyme and suggest that additional Cas9 orthologs may prove to exhibit similarly high accuracy, even without extensive engineering.

molecular biology

Genomics-Based Identification of Microorganisms in Human Ocular Body Fluid

Advances in genomics have the potential to revolutionize clinical diagnostics. Here, we examine the microbiome of vitreous (intraocular body fluid) from patients who developed endophthalmitis following cataract surgery or intravitreal injection. Endophthalmitis is an inflammation of the intraocular cavity and can lead to a permanent loss of vision. As controls, we included vitreous from endophthalmitis-negative patients, balanced salt solution used during vitrectomy, and DNA extraction blanks. We compared two DNA isolation procedures and found that an ultraclean production of reagents appeared to reduce background DNA in these low microbial biomass samples. We created a curated microbial genome database (>5700 genomes) and designed a metagenomics workflow with filtering steps to reduce DNA sequences originating from: i) human hosts, ii) ambiguousness/contaminants in public microbial reference genomes, and iii) the environment. Our metagenomic read classification revealed in nearly all cases the same microorganism than was determined in cultivation- and mass spectrometry-based analyses. For some patients, we identified the sequence type of the microorganism and antibiotic resistance genes through analyses of whole genome sequence (WGS) assemblies of isolates and metagenomic assemblies. Together, we conclude that genomics-based analyses of human ocular body fluid specimens can provide actionable information relevant to infectious disease management.

microbiology

Building a genome browser with GIVE

Growing popularity and diversity of genomic data demands portable and versatile genome browsers. Here, we present an open source programming library, called GIVE that facilitates creation of personalized genome browsers without requiring a system administrator. By inserting HTML tags, one can add to a personal webpage interactive visualization of multiple types of genomics data, including genome annotation, \"linear\" quantitative data (wiggle), and genome interaction data. GIVE includes a graphical interface called HUG (HTML Universal Generator) that automatically generates HTML code for displaying user chosen data, which can be copy-pasted into users personal website or saved and shared with collaborators. The simplicity of use was enabled by encapsulation of novel data communication and visualization technologies, including new data structures, a memory management method, and a double layer display method. GIVE is available at: http://www.givengine.org/.

bioinformatics

Limited role of differential fractionation in genome content variation and function in maize (Zea mays L.) inbred lines

Maize is a diverse paleotetraploid species with widespread presence/absence variation and copy number variation. One mechanism through which presence/absence variation can arise is differential fractionation. Fractionation refers to the loss of duplicate gene pairs from one of the maize subgenomes during diploidization and differential fractionation refers to non-shared gene loss events between individuals. We investigated the prevalence of presence/absence variation resulting from differential fractionation in the syntenic portion of the genome using two whole genome de novo assemblies of the inbred lines B73 and PH207. Between these two genomes, syntenic genes were highly conserved with less than 1% of syntenic genes being subject to differential fractionation. The few variable syntenic genes that were identified are unlikely to contribute to functional phenotypic variation, as there is a significant depletion of these genes in annotated gene sets. In further comparisons of 60 diverse inbred lines, non-syntenic genes were six times more likely to be variable compared to syntenic genes, suggesting that comparisons among additional genome assemblies are not likely to result in the discovery of large-scale presence/absence variation among syntenic genes.\n\nSIGNIFICANCE STATEMENTThere is a large amount of presence/absence variation for gene content in maize. One mechanism that has been hypothesized to contribute to this variation is differential fractionation between individuals following the maize whole genome duplication event. Using comparative genomics, with sorghum and rice representing the ancestral state, we observed little evidence of differential fractionation among elite inbred lines and the few differentially fractionated genes identified did not appear to confer functional significance.

plant biology

Identifying the genetic determinants of particular phenotypes in microbial genomes with very small training sets

Machine learning (ML) encompasses numerous algorithms that aim at discovering complex patterns between elements within large data using limited prior assumptions or modeling. However, some scientific disciplines still produce small data sets: in particular, empirical studies that try to find the mutations responsible for complex phenotypes are often limited to very small sample sizes (n), while scanning a large number of amino acid sites (p) in a proteome. To date, little is known on how ML performs in this type of so-called \"large p, small n\" problem. To address this question, we evaluated the performance of two general ML classifiers, adaptive boosting (AB) and random forest, on two data sets. To assess the impact of proteome size, we contrasted a small (viral) genome with a larger (bacterial) one. To analyze large proteomes, we further developed a chunking algorithm, and introduce a repeated random forest (RRF) algorithm that stabilizes model predictions. With the influenza data, we were able to rediscover amino acid sites experimentally implicated in three different complex phenotypes (infectivity, transmissibility, and pathogenicity). Results for the larger proteome, pertaining to three types of drug resistance (Ciprofloxacin, Ceftazidime, and Gentamicin), were more nuanced, with RRF making more sensible pre-dictions, with smaller errors rates, than AB. Furthermore, we show that chunking improved runtimes by an order of magnitude and may increase sensitivity of the predictions. Altogether, we demonstrate that ML algorithms can be used to identify genetic determinants in small proteomes (viruses), even with small numbers of individuals. We further show that even if the size of bacterial proteomes pushes AB to its limits in the context of small n, RRF may deserve more scrutiny, which should be facilitated by the plummeting costs of sequencing and, more critically, by phenotyping large cohorts of individuals.\n\nAuthor SummaryFinding the genetic determinants of a phenotype is typically performed by testing for an association between a particular allele and a trait, carrying out the testing over a large number of loci in a large cohort of individuals, itself divided into two subsets of individuals: those who have the trait (cases), and those who do not (controls). However, recruiting large cohorts can be problematic in some experimental fields, while using genotypic information rather than complete genomes can miss some mutations. To address these issues, we implemented two machine learning (ML) algorithms, tweaked for analyzing large genomes and providing stable results. The analysis of a small viral genome, for which genetic determinants of three phenotypes are already known, showed that our approach can rediscover known mutations, almost irrespective of the ML algorithm used. However, the analysis of a larger bacterial genome, for which genetic determinants of three phenotypes are unknown, suggested that the simpler of our modified algorithms performed better, returning more sensitive predictions with lower error rates. This work demonstrates the feasibility of finding genetic determinants of complex phenotypes based on a small number of complete genomes.

bioinformatics

Local and global chromatin interactions are altered by large genomic deletions associated with human brain development

BackgroundLarge copy number variants (CNVs) in the human genome are strongly associated with common neurodevelopmental, neuropsychiatric disorders such as schizophrenia and autism. Using Hi-C analysis of long-range chromosome interactions and ChIP-Seq analysis of regulatory histone marks we studied the epigenomic effects of the prominent large deletion CNV on chromosome 22q11.2 and also replicated a subset of the findings for the large deletion CNV on chromosome 1q21.1.\n\nResultsWe found that, in addition to local and global gene expression changes, there are pronounced and multilayered effects on chromatin states, chromosome folding and topological domains of the chromatin, that emanate from the large CNV locus. Regulatory histone marks are altered in the deletion proximal regions, and in opposing directions for activating and repressing marks. There are also significant changes of histone marks elsewhere along chromosome 22q and genome wide. Chromosome interaction patterns are weakened within the deletion boundaries and strengthened between the deletion proximal regions. We detected a change in the manner in which chromosome 22q folds onto itself, namely by increasing the long-range contacts between the telomeric end and the deletion proximal region. Further, the large CNV affects the topological domain that is spanning its genomic region. Finally, there is a widespread and complex effect on chromosome interactions genome-wide, i.e. involving all other autosomes, with some of the effect directly tied to the deletion region on 22q11.2.\n\nConclusionsThese findings suggest novel principles of how such large genomic deletions can alter nuclear organization and affect genomic molecular activity.

genetics