bioRxiv Science⌕ Search

Biology subjects

Khushi, K.

Publications and source records attributed to Khushi, K..

2 recordsLinked to original sources

The Chromosome-Scale Genome of Phyllanthus niruri Reveals Candidate Genes and a Putative Biosynthetic Framework for Phyllanthin Formation.

Phyllanthus niruri (Phyllanthaceae) is a medicinally important herb known for producing phyllanthin, a bioactive dibenzylbutane lignan with reported hepatoprotective and antioxidant properties. However, the biosynthetic basis of phyllanthin production remains unresolved, largely due to the absence of a reference genome for the species. We report a Chromosome-Scale Assembly of P. niruri generated by integrating PacBio HiFi long reads and Illumina short reads, followed by reference-guided scaffolding against Phyllanthus cochinchinensis. The assembly has an L50 of 7 and 97.6% BUSCO completeness. Annotation predicted 19,254 protein-coding genes (91.1% functionally annotated), with phenylpropanoid biosynthesis emerging as the most enriched specialized-metabolism pathway in the genome. Using pathway-guided genome mining, structural similarity analysis, and comparative metabolic reconstruction, we propose a putative biosynthetic pathway for phyllanthin originating from the phenylpropanoid-lignan branch through secoisolariciresinol-like intermediates, followed by terminal O-methylation reactions. A total of 305 unique candidate genes associated with the proposed pathway were identified, including expanded families of dirigent proteins, peroxidases, secoisolariciresinol dehydrogenases, and O-methyltransferases. Comparative transcriptomic analyses across related Phyllanthus species further supported the proposed pathway through coordinated expression of lignan-associated genes and tissue-specific enrichment of O-methyltransferases. This work provides the first reference genome for P. niruri and a prioritized candidate gene set for functional characterization of phyllanthin biosynthesis.

genomics↗

A consensus genome sequence for the social amoeba Dictyostelium giganteum

The life cycle of the Dictyostelid amoebae is unusual in that it alternates between a free-living solitary phase and an aggregative social phase. We used six previously collected Dictyostelium giganteum strains from distinct ecological niches in the Mudumalai Nature Reserve, India. From them, we generated short Illumina reads and assembled a consensus genome, comprising nuclear and mitochondrial genomes, representative of all six strains. The nuclear assembly has an AT content of 75.76%, accounting for 38.52 Mb, and resolves into five chromosome-scale scaffolds that are consistent with the published karyotype. Its N50 is 3.01 Mb and L50 is 5. The BUSCO analysis shows 3.9% fragmentation and 92.1% completeness. Genome assembly and completeness are also validated using [~]5,500 genes from a publicly available transcriptome dataset (PRJNA48443) derived from the post-aggregation stage. 13,251 predicted proteins are encoded by the genome, including ABC transporters, polyketide synthases, Ras/Rho GTPases, and expanded families of protein kinases. Comparative analysis demonstrates extensive conservation of syntenic blocks related to the dictyostelids D. discoideum and D. firmibasis as well as lineage-specific rearrangements. About 18% of the nuclear genome is made up of repetitive DNA, mostly in the form of simple repeats. Major transposable element classes, including piggyBac-like fragments, were found by homology searches. Long poly-asparagine/glutamine tracts are less common than in D. discoideum, but low-complexity sequences are common due to strong AT-driven codon bias. Comparative proteome-level orthology analysis across Dictyostelium species and Entamoeba identified a conserved Amoebozoan core together with a substantial Dictyostelium-specific gene repertoire. Domain-level comparisons further revealed widespread conservation of intracellular signalling and cytoskeletal modules shared with animals, whereas canonical metazoan extracellular adhesion domains were absent, highlighting the deep evolutionary roots of regulatory complexity underlying aggregative multicellularity.

genomics↗