bioRxiv Science⌕ Search

Biology subjects

Albert, J. L.

Publications and source records attributed to Albert, J. L..

2 recordsLinked to original sources

MrHAMER2: high-accuracy long-read RNA sequencing to decode isoform-specific variation in viral transcripts during latency

Alternative splicing (AS) greatly expands the repertoire of proteins encoded by the human genome. Viruses have been shown to hijack AS cellular pathways to sustain replication or lead to latency. In HIV-1 infection, the virus integrates into the host genome, becoming a transcriptional unit that directly engages in AS to regulate its gene expression. Sequencing advances have enabled insights into HIV-1 gene expression dynamics during productive replication. However, viral isoform dynamics during latency remain largely uncharacterized due to the low abundance of spliced viral transcripts in associated CD4+ T cell subsets, making their accurate detection and quantification challenging. MrHAMER2 is a high-accuracy long-read RNA sequencing method that leverages dual Unique Molecular Identifier (UMI) tagging of cDNA to accurately capture and quantify full-length isoforms with high dynamic range and 99.968% single-nucleotide accuracy. We used MrHAMER2 to decode the spliced HIV-1 transcriptome in a primary CD4+ T cell model of latency and showed substantial changes in viral isoforms bearing intron retentions accompanied by changes in their potential to generate translatable protein.

genomics↗

APHIX: Analysis Pipeline for HIV-1 Isoform eXploration Using Long-read RNA Sequencing Data

HIV-1 uses 4 major splice donors and 8 major splice acceptors as well as dozens of minor, cryptic, and uncharacterized splice sites to produce over one hundred distinct transcript isoforms from a single 9.2 kb genome. As a result, existing bioinformatic pipelines struggle to accurately analyze spliced HIV sequences due to the complex nature of HIV alternative splicing compared to human mRNA splicing. Previous approaches to identify HIV isoforms from long-read sequencing data used pipelines that are not publicly available, are convoluted to operate, or are locked into a specific HIV strain, which limits their wide adoption to other experimental designs or systems. To address this gap, we have developed a bioinformatic pipeline called APHIX that fully automates spliced isoform assignment, splice site usage quantification, and non-coding exon detection. APHIX takes a FASTQ/A of long-read transcripts and a HIV genome reference sequence and fully automates HIV isoform analysis. APHIX calculates splice site usage counts and percentages for each donor and acceptor site and their pairwise combinations, accurately assigns isoforms, and automatically identifies transcripts containing non-coding exons. APHIX is compatible with long-reads sequences generated from multiple platforms and library preps, including direct DNA and RNA sequencing. APHIX can also be adapted to multiple HIV-1 clades and strains by providing the appropriate reference sequence during bioinformatic processing. Overall, APHIX enables comprehensive processing of spliced sequences with reproducible results in a manner that is faster and easier to run compared to other methods.

bioinformatics↗