bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,675 records · Page 93Linked to original sources

Genome assembly of Bougainvillia cf. muscus (Cnidaria: Hydrozoa)

BackgroundAs one of just a handful of non-Bilaterian animal phyla, Cnidaria are key to understanding genome evolution across Metazoa. Despite their importance and diversity, the genomes of most species in the phylum are unsequenced, due in large part to difficulties cultivating them in a laboratory. Here, we present a genome sequence of Bougainvillia cf. mucus, a hydrozoan with four marginal bulbs each containing seven simple eyes (ocelli). This species appeared in our tanks from contamination. While we lacked sufficient samples for transcriptomic or functional studies, we were able to expand our knowledge of how the genome of this species compares to the few, better studied members of hydrozoans by investigating synteny to other cnidarians, repetitive element content, and phylogenetics and synteny of vision-related genes in this eyed species compared to eyeless relatives. ResultsThe genome sequence consists of 350 contigs with an N50 of 10 Mb, a total genome length of 375.328 Mb, a BUSCO score of 90.1%, and predicted protein coding genes totaling 46,431. We found a high degree of macrosynteny conservation with Hydra vulgaris and Hydractinia symbiolongicarpus. Repetitive elements make up 62% of this Bougainvillia genome. For vision-related genes, we identified 20 cnidarian opsins (cnidops) in Bougainvillia and found instances of gene duplication and loss in families associated with bilaterian eye development, phototransduction, and visual cycling. ConclusionsThis high-quality, contiguous genome in an eyed Hydrozoan will be a valuable resource for additional comparative genomic studies.

genomics↗

Targeted genomic surveillance of insecticide resistance in African malaria vectors

The emergence of insecticide resistance is threatening the efforts of malaria control programmes, which rely heavily on a limited arsenal of insecticidal tools, such as insecticide-treated bed nets. Importantly, genomic surveillance of malaria vectors can provide critical, policy-relevant insights into the presence and evolution of insecticide resistance, allowing us to maintain and extend the shelf life of these interventions. Yet the complex genetic architecture of resistance, combined with resource constraints in malaria-endemic settings, have thus far precluded the widespread use of genomics in routine surveillance. Meanwhile, stakeholders in sub-Saharan Africa are moving towards locally driven, decentralised generation of genomic data, underscoring the need for standardised and robust genomics workflows. To address this need, we demonstrate an approach to targeted genomic surveillance in Anopheles gambiae s.l with Illumina sequencing. We target 90 genomic loci in the Anopheles gambiae s.l genome, including 55 resistance-associated mutations and 35 ancestry informative markers. This protocol is coupled with advanced, automated software for accurate and reproducible variant analysis. We are able to elucidate population structure and ancestry in our cohorts and accurately identify most species in the An. gambiae species complex. We report frequencies of variants at insecticide-resistance loci and explore the continued evolution of the pyrethroid target site, the Voltage-gated sodium channel. Applying the platform to a recently established colony of field-caught resistant mosquitoes (Siaya, Kenya), we identified seven independent resistance-associated variants contributing to reduced efficacy of insecticide-treated nets in East Africa. Additionally, we leverage a machine learning algorithm (XGBoost) to demonstrate the possibility of predicting bioassay mortality using genotypes alone. This achieved high accuracy (75%), demonstrating the potential of targeted genomics to predictively monitor insecticide resistance. Together these tools provide a practical, scalable solution for resistance monitoring while advancing the goal of building local genomic surveillance capacity in sub-Saharan Africa.

genomics↗

A high-quality draft genome assembly of the Neotropical butterfly, Batesia hypochlora (Nymphalidae: Biblidinae)

We report a long-read high-coverage reference genome assembly of the Neotropical butterfly, Batesia hypochlora (Nymphalidae: Biblidinae). This represents the first reference genome in the Biblidinae subfamily, a clade subject to ongoing studies on seasonal and climate adaptation in the Amazon. We assembled the genome from PacBio HiFi long reads (66X coverage), polished it with Illumina short reads (15X coverage), and annotated it using PacBio IsoSeq RNA data. We observed 15 chromosome-sized scaffolds varying in length from 13.2 Mbp to 37.6 Mbp (median 24.3 Mbp), combining a total genome size of 395.788 Mbp. This assembly is highly contiguous (contig N50 of 25.14 Mbp) and complete (BUSCO completeness score of 98.6% and 0.2% duplication rate). Repeat annotation revealed that the genome consists of about one-third transposable elements. Gene prediction using RNAseq evidence uncovered 19,395 genes, of which 17,400 were assigned to 2,883 orthogroups when including genomes of the fruitfly, silk moth, and three other Nymphalid butterfly species. The high sequencing depth also allowed us to assemble the genomes of the mitochondria and the common endosymbiotic bacterium Wolbachia. The mitochondrial genome was fully assembled (15,540 bp in size) with all expected genes annotated. The Wolbachia genome was fragmented, and we determined that it belongs to the B-supergroup. The high-quality assembly of B. hypochlora can represent the subfamily in further comparative analysis of evolution and provide a key resource for ongoing work to explore reproductive biology and adaptations to seasonality in Amazonian butterflies.

genomics↗

Genomic and evolutionary factors influencing the prediction accuracy of optimal growth temperature in prokaryotes

Bacteria and archaea have evolved diverse genomic adaptations to thrive across various temperatures. These adaptations include genome sequence optimizations, such as increased GC content in rRNA and tRNA, shifts in codon and amino acids usage, and the acquisition of functional genes conferring adaptation for specific temperatures. Since the experimental determination of optimal growth temperatures (OGT) is only possible for cultured species, predicting OGT from genomic information has become increasingly important given the exponential increase in genomic data. Although previous studies developed prediction models integrating multiple features based on genome composition using machine learning, the accuracy was variable depending on the target species, with models performing well for thermophiles but less accurately for psychrophiles. In this study, we curated the OGT and genomic data of 2,869 bacterial species to develop a novel prediction model incorporating features reflecting genomic adaptation toward lower temperatures. We found that species with rapid OGT shifts from their ancestors, including psychrophiles, showed less accuracy in genome composition-based models. Incorporating the gene presence/absence information associated with the rapid changes in OGT improved the prediction accuracy for psychrophiles. We also observed that OGT in archaea is phylogenetically more conserved than in bacteria, which may lead to the long-term optimization of the genome composition and explain high predictability of OGT in archaea. These findings highlight the importance of integrating long- and short-term evolutionary adaptations for phenotype prediction models.

genomics↗

Improved genome assembly of double haploid Prunus persica siblings Lovell 2D and Lovell 5D and the peach NLRome

Prunus persica (peach) has long served as a model fruit tree for studying phenological events. It has a relatively small genome and exhibits tremendous plasticity in climate tolerances due to the high variation of chill requirements, bloom times, and fruit ripening times. The peach variety Lovell 2D was used to generate one of the first high-quality genome assemblies for a tree species, using Sanger sequencing of genetic-map ordered BAC clones. A key to the high quality of this early assembly was the use of a doubled haploid variety, which eliminates the challenges posed by mixed haplotypes. Here, we re-sequenced and assembled the Lovell 2D genome along with a doubled haploid sibling Lovell 5D using 3rd generation technologies. The resulting genomes were significantly more contiguous than the current Lovell 2D reference genome (ver2.0 updated in 2017) and are closer to the estimated total genome size for peach (265Mb). In addition, new gene, transposable element (TE), and Nucleotide-binding domain and Leucine-rich repeat receptor (NLR) annotations were performed to enhance the integrity and utility of the genome. These updated peach doubled-haploid reference assemblies will provide the research community with an improved reference genome for genomics-guided studies and breeding efforts.

genomics↗

GENERanno: A Genomic Foundation Model for Metagenomic Annotation

The rapid growth of genomic and metagenomic data has underscored the pressing need for advanced computational tools capable of deciphering complex biological sequences. In this study, we introduce Generanno, a compact yet powerful genomic foundation model (GFM) specifically optimized for metagenomic annotation. Trained on an extensive dataset comprising 715 billion base pairs (bp) of prokaryotic DNA, Generanno employs a transformer encoder architecture with 500 million parameters, enabling bidirectional attention over sequences up to 8192 bp at single-nucleotide resolution. This design addresses key limitations of existing methods, including the inability of traditional Hidden Markov Models (HMMs) to handle fragmented DNA sequences from multi-species microbial communities, as well as the suboptimal tokenization schemes of existing GFMs that compromise fine-grained analysis. At its core, Generanno excels in identifying coding regions from fragmented and mixed DNA sequences--a hallmark of metagenomic analysis. It achieves superior accuracy compared to traditional HMM-based methods (e.g., GLIMMER3, GeneMarkS2, Prodigal) and recent LLM-based approaches (e.g., GeneLM), while demonstrating robust generalization ability on archaeal genomes. Leveraging its advanced contextual understanding capability, Generanno further enables two essential functions: pseudogene prediction and taxonomic classification--both performed based solely on raw sequence data, without reliance on reference databases or comparative genomics. These functionalities collectively streamline the metagenomic analysis pipeline, significantly reducing preprocessing requirements and enabling end-to-end interpretation of sequencing data. Beyond its primary role in metagenomic annotation, Generanno also serves as a powerful GFM. To evaluate its broader utility, we curated the Prokaryotic Gener Tasks--a comprehensive benchmark suite specifically tailored for prokaryotic genomic analysis. It includes gene fitness prediction, antibiotic resistance identification, gene classification, and taxonomic classification, reflecting diverse aspects of functional genomics. On this benchmark, Generanno consistently outperforms existing GFMs such as DNABERT-2, NT-v2, and GenomeOcean, demonstrating strong generalization capabilities across a wide range of genomic tasks. Overall, Generanno provides a unified framework that integrates multiple critical functions for metagenomic annotation and beyond. By eliminating dependencies on external resources and offering rich contextual understanding of genomic sequences, this work delivers a foundational tool for advancing functional genomics in complex microbial communities. Implementation details and supplementary resources are available at https://github.com/GenerTeam/GENERanno.

genomics↗

Chromosome-level assembly and annotation of the grey reef shark (Carcharhinus amblyrhynchos) genome

To date only four of nine shark orders have nuclear reference genomes, despite next-generation sequencing advances. Particularly for threatened shark species, there is a lack of reliable genomes which are crucial in facilitating research and conservation applications. We assembled the first nuclear reference genome of the endangered grey reef shark (Carcharhinus amblyrhynchos) using long-read PacBio HiFi and Omni-C sequencing to reach chromosome-level contiguity (36 pseudo chromosomes; 2.9 Gbp) and high completeness (94% complete BUSCOs). BRAKER3 annotated 16,522 protein-coding genes after masking repetitive elements which accounted for 59% of the genome. We identified potential X and Y sex chromosomes on pseudo chromosomes 36 and 57, respectively. The quality and completeness of the draft genome of C. amblyrhynchos suggest that it has the potential to facilitate comprehensive comparative genomics, enabling researchers to investigate genetic variations and adaptations specific to this population and will help advance conservation genetic applications. Significance StatementStemming from an ancient vertebrate lineage, sharks present an interesting evolutionary study system. A third of shark species face extinction, yet critical genomic resources necessary for research and conservation remain scarce. To address this gap, we assembled and annotated the first chromosome-level nuclear reference genome of the threatened grey reef shark (Carcharhinus amblyrhynchos) at high completeness. This genome will help advance studies in evolution, phylogenetics, adaptation, and conservation, offering insights not only for this species but for wider elasmobranch and vertebrate research.

genomics↗

Mitochondrial Genome-Based Phylogeny of Turbellarians and Evidence for Accelerated Mitochondrial Evolution in Symbiotic Species

BackgroundFlatworms are a highly diverse phylum with over 26,500 predominantly parasitic species. A minor portion of this diversity comprise predominantly free-living "turbellarians" Phylogenetic relationships within turbellarian orders remain debated, with recent mitochondrial genome studies also questioning the monophyly of the "Neoophora clade". Some unique mitochondrial gene features have also been observed in this group. Within Turbellaria, the order Rhabdocoela includes significant symbiotic lineages, such as the endosymbiotic Umagillidae, Pterastericolidae, and Graffillidae, and the ectosymbiotic Temnocephalidae, which notably exhibits characteristics akin to a parasitic lifestyle. Given evidence linking parasitic lifestyles to accelerated mitogenomic evolution, we hypothesize that similar patterns: symbiotic turbellarians have mitochondrial genomes that accelerate evolution compared to free-living turbellarians. This study presents the first complete mitochondrial genome of the ectosymbiont Craspedella pedum, provides mitogenomic insights into turbellarian phylogeny, and addressing our hypothesis. ResultsThe mitochondrial genome of Craspedella pedum is a circular DNA molecule of 18,456 base pairs, containing the standard 36 flatworm mitochondrial genes, a duplicated trnT, two cox1 pseudogene fragments, a putative atp8 gene, and several distinctive NCRs. Phylogenetic analyses based on 47 mitochondrial genomes, including the newly sequenced C. pedum and two assembled species from SRA database, using CAT-GTR model and BI and ML algorithms further confirmed that the paraphyly of the Neoophora clade and the basal position of Catenulida and Macrostomida. Different from previous study, Rhabdocoela forms a distinct clade within the turbellarians diverged immediately after Macrostomida before Polycladida and Tricladida. Within Rhabdocoela, C. pedum formed a direct sister clade with Typhloplanidae, suggesting a close phylogenetic relationship between the two. We also identified the paraphyletic of the Planoceridae of Polycladida and the unstable position of Planariidae within Tricladidain. Furthermore, we found a trend that the nad4L-nad4 gene box in turbellarians may have evolved from an overlapping state, to the insertion of a non-coding region (NCR), and subsequently to a separated configuration, correlating with species divergence. The selection pressure analysis showcased selective relaxation from free-living species of Rhabdocoela to symbiotic species of Rhabdocoela. Furthermore, we also detected a relaxation from certain tubellarian lineages to Rhabdocoela. Along with higher GORR and longer Brl, Rhabdocoela and its symbiotic group possesses a more rapidly evolving mitochondrial genome. ConclusionsThe mitochondrial genome of Craspedella pedum displays uncommon characteristics, combined extended branch lengths and elevated GORR, suggesting rapid evolution and extensive rearrangements. Our phylogenetic analysis, integrating additional mitogenome data, further corroborates the paraphyly of the Neoophora clade and offers novel insights into turbellarian phylogeny. For the first time, we confirmed the accelerated evolution of mitochondrial genomes in symbiotic tubellarians compared to free-living ones, as evidenced by a longer average Brl, a higher average GORR, and relaxed selection pressure. Additionally, with the same evidences, we found that Rhabdocoela also exhibited accelerated mitochondrial genome evolution in planarians.

genomics↗

Programmed DNA elimination drives rapid genomic innovation in two thirds of all bird species

Bird genomes are among the most stable in terms of synteny and gene content across vertebrates. However, germline-restricted chromosomes (GRCs) represent a striking exception where programmed DNA elimination confines large-scale genomic changes to the germline. GRCs are known to occur in songbirds (oscines), but have been studied only in a few species of Passerides such as the zebra finch, the key model for passerine genomics. Their presence and evolutionary dynamics in most major passerine lineages remain largely unexplored, with suboscines entirely unexamined by cytogenetic or genomic methods. Here, we present the most comprehensive comparative analysis of GRCs to date, spanning 44 million years of passerine evolution. By generating the first germline reference genomes of an oscine and a suboscine, 22 novel germline draft genomes spanning nearly all major passerine lineages and a germline draft genome of a parrot outgroup, we show that the GRC is likely present in 6,700 passerine species. Our results reveal that the GRC evolves rapidly and distinctly from the standard A chromosomes (autosomes and sex chromosomes), yet retains functionally important, selectively maintained genes. We observed gene and repeat turnover occuring orders of magnitude faster than on the A chromosomes. Some GRC genes, such as cpeb1 and pim1, are widespread from an ancient duplication. In contrast, other GRC genes, like mfsd2b and bmp15, have been independently duplicated onto the GRC multiple times, suggesting adaptive constraints. The discovery of zglp1 on the zebra finch GRC, initially copied from chromosome 30 and subsequently lost from it, indicates functional replacement, where the GRC permits gene loss from the standard genome. As the GRC harbors the only zglp1 copy in most of the [~]4000 Passerides species, GRC loss would compromise essential germline functions. Our findings establish the GRC as a genomic innovator driving rapid germline evolution. This fact highlights its evolutionary significance for passerine diversification and suggests that programmed DNA elimination may be an overlooked yet phylogenetically widespread mechanism in many understudied animal lineages.

genomics↗

The reference genome for the northeastern Pacific bull kelp, Nereocystis luetkeana

Bull kelp, Nereocystis luetkeana, is a northeastern Pacific kelp with broad distribution from Alaska to central California. Its population declines have caused severe concerns in northern California, the Salish Sea in Washington, and recently in some populations in Oregon. Despite bull kelps accumulated ecological and physiological studies, an assembled and annotated genomic reference was still unavailable. Here, we report the complete and annotated genome of Nereocystis luetkeana, produced by the California Conservation Genomics Project (CCGP), which aims to reveal genomic diversity patterns across California by sequencing the complete genomes of approximately 150 carefully selected species. The genome was assembled into 1,562 scaffolds with 449.82 Mb, 80x of coverage and 22,952 gene models. BUSCO assembly showed a completeness score of 72% for the stramenopiles gene set. The mitochondria and chloroplast genome sequences have 37 Mb and 131 Mb, respectively. The orthology analysis between 10 Phaeophycean genomes showed 1,065 expanded and 286 unique orthogroups for this species. Pairwise comparisons showed 542 orthogroups present only in N. luetkeana and M. pyrifera, another large-body kelp. The enrichment analysis of these orthogroups showed important functions related to central metabolism and signaling due to ATPases enrichment in these two species. This genome assembly will provide an essential resource for the ecology, evolution, conservation, and breeding of bull kelp.

genomics↗

An updated resource of 180K soybean SNP genotyping array based on the T2T reference genome

Single nucleotide polymorphism (SNP) genotyping has revolutionized crop improvement by enabling high-resolution genomic analyses and accelerating breeding programs. In soybean (Glycine max (L.) Merr.), a globally important legume crop, existing genotyping data for 180,961 SNP markers from the Korean soybean core collection were generated using the outdated Williams 82 reference genome version 1 (Wm82.v1), which contains numerous assembly gaps and misassemblies that limit genomic resolution. While high-quality reference genomes including Wm82.v4 and Wm82.v6 (telomere-to-telomere assembly) are now available, the valuable existing SNP array data have not been integrated with these improved genomic resources. Here we show successful remapping of the 180K SNP array data to both Wm82.v4 and Wm82.v6 reference genomes through sequence-based alignment of flanking regions. We extracted flanking sequences from SNP marker positions in Wm82.v1 and mapped them to the newer reference versions based on sequence similarity, excluding markers with mapping failures, allele mismatches, low-identity alignments, or multiple mappings, which resulted in successful mapping of 175,202 and 175,763 markers to Wm82.v4 and Wm82.v6, respectively. We also remapped genotype data from 927 soybean accessions (497 USDA-GRIN accessions and 430 Korean core collection accessions) to both reference versions. This updated SNP dataset provides the soybean research community with a comprehensive genomic resource that leverages both existing genotyping investments and state-of-the-art reference genome assemblies for enhanced crop improvement and genomic studies.

genomics↗

Integrative genomics identifies candidate genes underlying trypanotolerance in hybrid African cattle

Integrative genomics combines data from different omic sources to link genotypes and phenotypes with the aim of unravelling biological networks and pathways that undergird complex traits, particularly with respect to disease. In this respect, integrative genomics using population and functional genomic data can be employed to understand evolutionary processes that have shaped adaptation to infectious diseases in domestic cattle. This approach can be particularly informative for African cattle, which exhibit a complex mosaic of Bos taurus (taurine) and Bos indicus (indicine) genomic ancestry. Some African taurine populations have an important evolutionary adaptation known as trypanotolerance, a genetically determined tolerance of infection by trypanosome parasites (Trypanosoma spp.) that cause African animal trypanosomiasis (AAT) disease. AAT is one of the largest constraints to livestock production in sub-Saharan Africa and causes a financial burden of approximately $4.5 billion annually. In this study we identified putative candidate genes underlying trypanotolerance through the integration of local ancestry inference (LAI) from genome-wide SNP data for multiple trypanotolerant and trypanosusceptible hybrid cattle populations with RNA-seq and expression microarray transcriptomics data from multiple tissues collected across time course trypanosome infection experiments. These candidate genes included AGO2, CBL, CNOT1, EDN1, IL1B, NFKB1, RIPK1, and TRAF2. Functional analysis of the gene set outputs from this work highlighted GO terms associated with the immune system (including the major histocompatibility complex - MHC) and cell signalling processes. These results signpost future work to elucidate the cellular networks and pathways that drive trypanotolerance. Author SummaryIntegrative genomics combines different types of data to identify links between genes and traits, particularly with respect to disease. In this respect, integrative genomics can be used to understand the admixture and adaptation to infectious diseases that have shaped the genomes of domestic cattle. This is particularly noticeable in the case of African cattle, which form a complex mosaic of Bos taurus (taurine) and Bos indicus (indicine) ancestry. Some African taurine populations exhibit an evolutionary adaptation known as trypanotolerance, a genetically determined tolerance of infection by trypanosome parasites (Trypanosoma spp.) that cause African animal trypanosomiasis (AAT) disease. AAT is one of the largest constraints to livestock production in sub-Saharan Africa and causes a financial burden of approximately $4.5 billion annually. In this study we identify potential candidate genes underlying trypanotolerance through integration of subchromosomal genomic ancestry data from multiple trypanotolerant and trypanosusceptible hybrid cattle populations with gene expression data from multiple tissues collected across time course trypanosome infection experiments.

genomics↗

Chromosome-scale genome assembly of Malcolmia littorea using long-read sequencing and single-pollen genotyping technologies

Malcolmia littorea, a member of the family Brassicaceae, is adapted to coastal and sandy environments and has become a model in studies of reproductive barriers. However, genomic resources for the species are limited. Here, with the aim of understanding the molecular mechanisms underlying key traits in M. littorea, including its survival under harsh conditions, we present a de novo genome assembly consisting of 10 chromosome-scale sequences. We employed a high-fidelity long-read sequencing technology for genome assembly. To anchor the sequences to chromosomes, we developed a single-pollen genotyping method to construct a genetic linkage map based on SNPs derived from transcriptomes of pollen grains, possessing recombinant haploid genomes. We built a genome assembly consisting of 10 chromosome-scale sequences (214 Mb in total) for M. littorea containing 30,861 predicted genes. A comparative genome analysis and gene prediction indicated that the genome of M. littorea is double the size of the Arabidopsis thaliana genome, consistent with a whole-genome duplication followed by gene subfunctionalization and/or neofunctionalization in M. littorea. This study provides a basis for research on M. littorea, an understudied species with ecological and evolutionary significance.

genomics↗

First Genome-Wide Centromere Map of Trypanosoma cruzi Reveals Linear and 3D Compartment Boundaries and Spatial Clustering

Background: Trypanosoma cruzi, the etiological agent of Chagas disease, possesses a highly repetitive genome that has historically hindered high-quality assembly and structural characterization. Despite significant advances in assembling T. cruzi genomes, major gaps remain. Among these, the complete repertoire of centromeric sequences has remained elusive, representing a critical missing piece in our understanding of chromosome structure and inheritance. Results: Here, we generated high-coverage Hi-C (genome-wide chromosome conformation capture) data for the widely used T. cruzi Dm28c strain improving its genome assembly, reducing the number of scaffolds and producing a more contiguous and accurate genome. To investigate centromere organization, we performed ChIP-seq using the mNeonGreen-myc-tagged kinetochore proteins KKT2 and KKT3, resulting in the identification of 40 KKT-enriched peaks across 29 scaffolds. These peaks were located in regions enriched in retrotransposable elements, particularly L1Tc and VIPER, near strand switch regions, areas of high GC content, and at the boundaries between conserved genes and virulence-factor multigene families. Conclusion: Notably, Hi-C analysis revealed that centromeres may act as structural boundaries contributing to genome compartmentalization and frequently engage in 3D spatial clustering, suggesting a role in higher-order nuclear architecture. Overall, our study provides a high-quality reference genome for the Dm28c strain, presents the first genome-wide centromere map in T. cruzi, and offers novel insights into centromere-mediated 3D genome organization

genomics↗

The Drosophila OSC Genome: A Resource for Studies of Transposon and piRNA Biology

Accurate genome assemblies are critical for understanding small RNA-mediated genome defense. In animals, the PIWI-interacting RNA (piRNA) pathway protects genome integrity by silencing transposable elements. Studying how piRNAs are generated and how they guide heterochromatin formation requires complete reconstruction of genomic piRNA source loci and detailed transposon maps. Here, we present a high-quality de novo genome assembly of Drosophila melanogaster ovarian somatic cells (OSCs), a widely used cell line that recapitulates nuclear piRNA biology. The OSC genome differs substantially from the reference genome, with major differences in transposon content and piRNA cluster composition. Our assembly resolves the 700 kb flamenco locus, the primary piRNA cluster in OSCs, and provides a genome-wide transposon map. Using this resource, we characterize piRNA source loci, reveal how piRNA cluster composition determines transposon-derived piRNA profiles, and clarify the widespread impact of the nuclear piRNA pathway on heterochromatin. Finally, we provide an open platform for integrating user-generated datasets with the OSC genome, creating a community resource for studying transposon control and piRNA biology.

genomics↗

Tick Genome Assemblies: Overcoming biological limitations through advances in sequencing and assembly

Ticks are blood-feeding arthropods with approximately 1,000 species, however, only 24 species currently have a genome assembly. These genome assemblies are important resources to advance tick biology and control of tick-associated diseases. Generating tick genome assemblies is challenging due to their small body size (low DNA input), DNA contamination (from microbiota and host bloodmeals), large genome size (on average 2.4 Gbp for hard ticks), and abundant transposable elements (at least 68% of the assembly for Ixodes species). Advances in sequencing technologies have driven an increasing number of tick assemblies from 2011 to 2025. We characterize and assess the 54 tick genome assemblies within public genome databases using QUAST-LG and BUSCO compleasm. Then we evaluate the impact of biological source material and sequencing platforms on these tick genome assemblies. From the 54 tick assemblies, we identify 34 high-quality assemblies from 21 species that are suitable for downstream analyses. We recommend future tick genome assemblies use long-read sequencing platforms and Hi-C scaffolding to improve genomic resources for these unique blood-feeding parasites.

genomics↗

Do DNA and cytometric measures agree on genome sizes?

Measurement of DNA contents of genomes is valuable for understanding genome biology, including assessments of genome assemblies, but it is not a trivial problem. Measuring contents of DNA shotgun reads is complicated by several factors: biological contents of genomes, laboratory methods, sequencing technology and computational processing. This compares and shares complications with cytometric measures of genome size and contents. There is an obvious discrepancy between cytometry and current long-read assemblies: assemblies average significantly below cytometric sizes. Measures of population changes within a species will control some of these complications. This report examines five species population sets with published cytometric and DNA data sets: Arabidopsis thaliana and A. arenosa, Arctic plants of Cochlearia genus, clonal populations of a rotifer Brachionus asplanchnoidis, and Zea mays corn plants. Results of this are clear, if complicated: DNA and cytometry do measure the same genome sizes, when done carefully with controls or adjustments for errors. Population changes in genome sizes are found by assembly-mapped measures of DNA, including environment or regional effects of latitude and altitude. Copy numbers of repeats, transposons and genes are changing. Kmer-based measures of DNA generally fail to match cytometry, miss population changes, and are opaque to understanding measurement errors. Oxford Nanopore technology produces the least biased DNA for measurement, with recent ONT.R10 Simplex data a match to cytometric sizes for corn, tomato plants, zebra fish and zebra finch bird. Assemblies of these species DNA average 12% below measured sizes, incomplete for duplicated content. New assembly of this ONT Simplex DNA reaches the size measured by cytometry and Gnodes, in 4 of 5 species. rRNA gene duplications measure one aspect of this discrepancy: the genome assemblies examined all are missing many rRNA genes. New experiments that measure both cytometry and DNA, controlling error factors, are warranted to clarify these results and suggest improvements for genome projects.

genomics↗

Resolving the Full Spectrum of Human Genome Variation using Linked-Reads

Large-scale population based analyses coupled with advances in technology have demonstrated that the human genome is more diverse than originally thought. To date, this diversity has largely been uncovered using short read whole genome sequencing. However, standard short-read approaches, used primarily due to accuracy, throughput and costs, fail to give a complete picture of a genome. They struggle to identify large, balanced structural events, cannot access repetitive regions of the genome and fail to resolve the human genome into its two haplotypes. Here we describe an approach that retains long range information while harnessing the advantages of short reads. Starting from only [~]1ng of DNA, we produce barcoded short read libraries. The use of novel informatic approaches allows for the barcoded short reads to be associated with the long molecules of origin producing a novel datatype known as Linked-Reads. This approach allows for simultaneous detection of small and large variants from a single Linked-Read library. We have previously demonstrated the utility of whole genome Linked-Reads (lrWGS) for performing diploid, de novo assembly of individual genomes (Weisenfeld et al. 2017). In this manuscript, we show the advantages of Linked-Reads over standard short read approaches for reference based analysis. We demonstrate the ability of Linked-Reads to reconstruct megabase scale haplotypes and to recover parts of the genome that are typically inaccessible to short reads, including phenotypically important genes such as STRC, SMN1 and SMN2. We demonstrate the ability of both lrWGS and Linked-Read Whole Exome Sequencing (lrWES) to identify complex structural variations, including balanced events, single exon deletions, and single exon duplications. The data presented here show that Linked-Reads provide a scalable approach for comprehensive genome analysis that is not possible using short reads alone.

genomics↗