Benchmarking long-read RNA sequencing for de novo transcriptome assembly in non-model plant species: insights from Moricandia arvensis
De novo transcriptome assembly is the standard approach for constructing a reference transcriptome in non-model plants that lack a high-quality genome, yet short-read assemblies struggle to resolve full-length isoforms. Long-read Iso-Seq (PacBio) captures full-length transcripts directly, but its use as a primary reference and the choice of downstream assembly pipeline remains poorly benchmarked. Here, we construct genome-free Iso-Seq reference transcriptomes for two organs, flower and leaf, of the non-model species Moricandia arvensis (L.) DC. (Brassicaceae), and systematically compare pipeline strategies combining Iso-Seq clustering, CD-HIT redundancy reduction, and Cogent graph-based reconstruction, benchmarked by BUSCO completeness, RSEM short-read mapping, and TransDecoder ORF completeness. We find that the optimal pipeline is organ specific. For flower, CD-HIT pre-filtering followed by Cogent reconstruction produced a high-quality reference (95.3% BUSCO complete). For the leaf, the same Cogent step was detrimental, reducing BUSCO completeness from 90.1% to 78.0% by incorrectly merging distinct genes; therefore, CD-HIT at 95% identity without reconstruction was retained. We trace this divergence to organ-specific input-data characteristics: leaf transcripts show extreme full-length-read expression skew and predominantly single-isoform gene support, depriving Cogent's graph algorithm of the multi-isoform evidence it requires. We find that the concentration of full-length reads among the most highly expressed transcripts predicts pipeline suitability before reconstruction, with per-transcript read depth acting as a necessary but non-discriminating floor. Because the leaf reference lacked gene-level structure, we further recovered gene-isoform grouping using expression-aware read-clustering (Corset), which preserved completeness while restoring the paralog structure expected of a paleopolyploid genome and outperformed sequence-only clustering. Our results provide a robust, genome-free framework for constructing full-length reference transcriptomes in non-model plant species and demonstrate that pipeline choice must be evaluated per organ rather than assuming one size fits all.