bioRxiv Science⌕ Search

Biology subjects

Scaketti, M.

Publications and source records attributed to Scaketti, M..

2 recordsLinked to original sources

First chromosome scale genome of Acrocomia aculeata

The Arecaceae family comprises economically, and ecologically significant palm species widely distributed across tropical and subtropical regions. Among them, Acrocomia aculeata, commonly known as Macauba, has gained attention due to its high oil yield, environmental adaptability, and potential applications in sustainable agriculture and bioenergy production. This study presents the first chromosome scale genome assembly of A. aculeata and the first reference genome of the genus, using Oxford Nanopore Technologies sequencing, PacBio HiFi and Hi-C proximity ligation data. The genome covers 1.94 Gbp in the 15 pseudochromosomes, with N50 of 143.43 Mbp, achieving a base quality (QV) of 76.4 and a k-mer completeness of 90.58% of PacBio HiFi data, with k=21. Interspersed repetitive regions represent 77.36% of the total genome, with long terminal repeat (LTRs) accounting for 54.98%. A total of 28.215 protein-coding genes were predicted. The genome was validated with BUSCO, achieving a completeness of 98.9% for predicted proteins. The first chromosome scale genome of Macauba provide new insights and breakthroughs for studies of conservation genomics and plant breeding programs.

genomics↗

Sample Size Impact (SaSii): an R script for estimating optimal sample sizes in population genetics and population genomics studies

Obtaining large sample sizes for genetic studies can be challenging, time-consuming, and expensive. However, small sample sizes may generate biased or imprecise results. Many studies have suggested the minimum sample size necessary to obtain robust and reliable results, but it is not possible to define one ideal minimum sample size that fits all studies. Here, we aim to present SaSii (Sample Size Impact), a R script to help researchers to define minimum sample size and to indicate minimum sample size patterns for some taxa groups. The patterns were obtained by analyzing previously published datasets with SaSii and can be used as a starting point for the sample design of population genetics and genomic studies. Our results showed that it is possible to estimate an adequate sample size that accurately represents the real population and does not require the scientist to take time to write any program code, extract and sequence samples or use population genetics programs, making it easier to gather this information. We also confirmed that sample sizes of five to twenty-five for SNP and fifteen to thirty for SSR can be used for most plant species, giving a better direction for new studies.

bioinformatics↗