bioRxiv ScienceSearch

Biology subjects

Kasif, S.

Publications and source records attributed to Kasif, S..

2 recordsLinked to original sources

Unexpected Properties of Short Genomic Tandem Repeats

Length polymorphisms in genomic short tandem repeats have been implicated in a variety of diseases, most notably human neurodegenerative disorders. Expansions of tandem repeats are also associated with genomic instability in cancer. Our previous study of length-3 tandem repeats uncovered a surprising pattern in the length distribution of certain such repeats in the non-coding regions of the human reference genome: a bias towards repeats of length 3n - 1, (n > 3). That is, the observed frequency of repeats of this length in the human genome is higher than expected by chance based on the frequency of shorter repeats.\n\nWe have hypothesized that this pattern may be a general property of genomic DNA. If true, this could have implications with regard to the dynamics of repeat expansion generally. To test this hypothesis, we have analyzed the genomic sequences of a broad range of eukaryotic organisms as well as several complete human genomes and obtained a number of thought provoking results. We establish that this unexpected elevation in frequency of 3n - 1 long repeats is statistically significant. We also expanded this analysis to different classes of genomic regions and tandem repeats of length four and five. The specific pattern was found in 13 of the 20 organisms analyzed, including all chordate and insect genomes tested. The bias pattern, however, was not confined to a single branch of the evolutionary tree. For some genomes, such as Drosophila melanogaster, the repeat bias surprisingly was also identified in exons. The pattern is present in both small and large genomes. A similar pattern was also found in tetranucleotide and pentanucleotide repeats in the human genome. Another surprising property was identified for the flanking GC content for triplet repeats of length 3n. These findings indicate a puzzling new genomic phenomenon with possible evolutionary and disease-related implications.

bioinformatics

Not All Experimental Questions Are Created Equal: Accelerating Biological Data to Knowledge Transformation (BD2K) via Science Informatics, Active Learning and Artificial Intelligence

Pablo Picasso, when first told about computers, famously quipped \"Computers are useless. They can only give you answers.\" Indeed, the majority of effort in the first half-century of computational research has focused on methods for producing answers. Incredible progress has been achieved in computational modeling, simulation and optimization, across domains as diverse as astrophysics, climate studies, biomedicine, architecture, and chess. However, the use of computers to pose new questions, or prioritize existing ones, has thus far been quite limited.\n\nPicassos comment highlights the point that good questions can sometimes be more elusive than good answers. The history of science offers numerous examples of the impact of good questions. Paul Erd[o]s, the wandering monk of mathematical graph theory, offered small prizes for anyone who could prove conjectures he identified as important (1). The prizes varied in cash amounts based on the perceived complexity of the problem posed by Erd[o]s.\n\nPosing technical questions and allocating resources to answer them has taken on a new guise in the Internet age. The X-Prize foundation (http://www.xprize.org/) offers multi-million dollar bounties for grand technological goals, including goals for sequencing genomes or space exploration. Several companies provide portals where customers can place cash bounties on educational, scientific or technological challenges, while potential problem solvers can compete to produce the best solutions for these problems. Amazons Turk site (https://www.mturk.com/mturk/welcome) links people requesting performance of intellectual tasks to people willing to work on them for a fee. Such crowd-sourcing systems create markets of questions and answers, and can help allocate resources and capabilities efficiently.\n\nThis paradigm suggests a number of interesting questions for scientific research. In a resource limited environment, can funds and research capacity be allocated more efficiently? Can knowledge demand provide an alternative or complementary mechanism to traditional investigator-initiated research grants?\n\nThe fathers of Artificial Intelligence (AI) and Herbert Simon in particular envisioned the application of AI to Scientific Discovery in different forms and styles (focusing on physics). We follow on these early dreams and describe a novel approach aimed at remodeling of the biomedical research infrastructure and catalyze gene function determination. We aim to start a bold discussion of new ideas aimed towards increasing the efficiency of the allocation of research capacities, reproducibility, provenance tracking, removing redundancy and catalyzing knowledge gain with each experiment. In particular, we describe a tractable computational framework and infrastructure that can help researchers assess the potential information gain of millions of experiments before conducting them. The utility of experiments in this case is modeled as the predictive knowledge (formalized as information) to be gained as a result of performing the experiment. The experimentalist would then be empowered to select experiments that maximized information gain if they wished, recognizing that there are frequently other considerations, such as a specific technological or medical utility, that might over-ride the priority of maximizing information gain. The conceptual approach we develop is general, and here we apply it to the study of gene function.

bioinformatics