bioRxiv ScienceSearch

Biology subjects

Lunter, G.

Publications and source records attributed to Lunter, G..

3 recordsLinked to original sources

Demographic inference using particle filters for continuous Markov jump processes

Demographic events shape a populations genetic diversity, a process described by the coalescent-with-recombination (CwR) model that relates demography and genetics by an unobserved sequence of genealogies. The space of genealogies over genomes is large and complex, making inference under this model challenging. We approximate the CwR with a continuous-time and -space Markov jump process. We develop a particle filter for such processes, using way-points to reduce the problem to the discrete-time case, and generalising the Auxiliary Particle Filter for discrete-time models. We use Variational Bayes for parameter inference to model the uncertainty in parameter estimates for rare events, avoiding biases seen with Expectation Maximization. Using real and simulated genomes, we show that past population sizes can be accurately inferred over a larger range of epochs than was previously possible, opening the possibility of jointly analyzing multiple genomes under complex demographic models. Code is available at https://github.com/luntergroup/smcsmc MSC 2010 subject classificationsPrimary 60G55, 62M05, 62M20, 62F15; secondary 92D25.

evolutionary biology

An Equivariant Bayesian Convolutional Network predicts recombination hotspots and accurately resolves binding motifs

MotivationConvolutional neural networks (CNNs) have been trememdously successful in many contexts, particularly where training data is abundant and signal-to-noise ratios are large. However, when predicting noisily observed biological phenotypes from DNA sequence, each training instance is only weakly informative, and the amount of training data is often fundamentally limited, emphasizing the need for methods that make optimal use of training data and any structure inherent in the model.\n\nResultsHere we show how to combine equivariant networks, a general mathematical framework for handling exact symmetries in CNNs, with Bayesian dropout, a version of MC dropout suggested by a reinterpretation of dropout as a variational Bayesian approximation, to develop a model that exhibits exact reverse-complement symmetry and is more resistant to overtraining. We find that this model has increased power and generalizability, resulting in significantly better predictive accuracy compared to standard CNN implementations and state-of-art deep-learning-based motif finders. We use our network to predict recombination hotspots from sequence, and identify high-resolution binding motifs for the recombination-initiation protein PRDM9, which were recently validated by high-resolution assays. The network achieves a predictive accuracy comparable to that attainable by a direct assay of the H3K4me3 histone mark, a proxy for PRDM9 binding.\n\nAvailabilityhttps://github.com/luntergroup/EquivariantNetworks\n\nContactrichard.brown@well.ox.ac.uk, gerton.lunter@well.ox.ac.uk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

genomics

A high throughput screen for active human transposable elements

Transposable elements (TEs) are mobile genetic sequences that randomly propagate within their hosts genome. This mobility has the potential to affect gene transcription and cause disease. However, TEs are technically challenging to identify, which complicates efforts to assess the impact of TE insertions on disease. Here we present a targeted sequencing protocol and computational pipeline to identify polymorphic and novel TE insertions using next-generation sequencing: TE-NGS. The method simultaneously targets the three subfamilies that are responsible for the majority of recent TE activity (L1HS, AluYa5/8, and AluYb8/9) thereby obviating the need for multiple experiments and reducing the amount of input material required. Here we describe the laboratory protocol and detection algorithm, and a benchmark experiment for the reference genome NA12878. We demonstrate a substantial enrichment for on-target fragments, and high sensitivity and precision to both reference and NA12878-specific insertions. We report 17 previously unreported loci for this individual which are supported by orthogonal long-read evidence, and we identify 1,471 polymorphic and novel TEs in 12 additional samples that were previously undocumented in databases of insertion polymorphisms. We anticipate that future applications of TE-NGS alongside exome sequencing of patients with sporadic disease will reduce the number of unresolved cases, and improve estimates of the contribution of TEs to human genetic disease.

genomics