bioRxiv Science⌕ Search

Biology subjects

Cappio Barazzone, E.

Publications and source records attributed to Cappio Barazzone, E..

2 recordsLinked to original sources

ultraID: a compact and efficient enzyme for proximity-dependent biotinylation in living cells

Proximity-dependent biotinylation (PDB) combined with mass spectrometry analysis has established itself as a key technology to study protein-protein interactions in living cells. A widespread approach, BioID, uses an abortive variant of the E. coli BirA biotin protein ligase, a quite bulky enzyme with slow labeling kinetics. To improve PDB versatility and speed, various enzymes have been developed by different approaches. Here we present a novel small-size engineered enzyme: ultraID. We show its practical use to probe the interactome of Argonaute-2 after a 10 min labeling pulse and expression at physiological levels. Moreover, using ultraID, we provide a membrane-associated interactome of coatomer, the coat protein complex of COPI vesicles. To date, ultraID is the smallest and most efficient biotin ligase available for PDB and offers the possibility of investigating interactomes at a high temporal resolution.

biochemistry↗

Accurate de novo identification of biosynthetic gene clusters with GECCO

Biosynthetic gene clusters (BGCs) are enticing targets for (meta)genomic mining efforts, as they may encode novel, specialized metabolites with potential uses in medicine and biotechnology. Here, we describe GECCO (GEne Cluster prediction with COnditional random fields; https://gecco.embl.de), a high-precision, scalable method for identifying novel BGCs in (meta)genomic data using conditional random fields (CRFs). Based on an extensive evaluation of de novo BGC prediction, we found GECCO to be more accurate and over 3x faster than a state-of-the-art deep learning approach. When applied to over 12,000 genomes, GECCO identified nearly twice as many BGCs compared to a rule-based approach, while achieving higher accuracy than other machine learning approaches. Introspection of the GECCO CRF revealed that its predictions rely on protein domains with both known and novel associations to secondary metabolism. The method developed here represents a scalable, interpretable machine learning approach, which can identify BGCs de novo with high precision.

bioinformatics↗