Hostname: page-component-76d6cb85b7-dqfph Total loading time: 0 Render date: 2026-07-20T01:43:36.983Z Has data issue: false hasContentIssue false

Computational Identification of Repeat-Containing Proteins and Systems

Published online by Cambridge University Press:  20 October 2020

Han Altae-Tran
Affiliation:
Broad Institute of MIT and Harvard Cambridge, Cambridge, MA 02142, USA Department of Biological Engineering, Massachusetts Institute of Technology, Cambridge, MA 02139, USA
Linyi Gao
Affiliation:
Broad Institute of MIT and Harvard Cambridge, Cambridge, MA 02142, USA Department of Biological Engineering, Massachusetts Institute of Technology, Cambridge, MA 02139, USA
Jonathan Strecker
Affiliation:
Broad Institute of MIT and Harvard Cambridge, Cambridge, MA 02142, USA
Rhiannon K. Macrae
Affiliation:
Broad Institute of MIT and Harvard Cambridge, Cambridge, MA 02142, USA McGovern Institute for Brain Research, Massachusetts Institute of Technology, Cambridge, MA 02139, USA
Feng Zhang*
Affiliation:
Broad Institute of MIT and Harvard Cambridge, Cambridge, MA 02142, USA Department of Biological Engineering, Massachusetts Institute of Technology, Cambridge, MA 02139, USA Department of Brain and Cognitive Sciences, Massachusetts Institute of Technology, Cambridge, MA 02139, USA McGovern Institute for Brain Research, Massachusetts Institute of Technology, Cambridge, MA 02139, USA Howard Hughes Medical Institute, Cambridge, MA 02139, USA
*
Author for correspondence: *Correspondence to: Feng Zhang, E-mail: zhang@broadinstitute.org
Rights & Permissions [Opens in a new window]

Abstract

Repetitive sequence elements in proteins and nucleic acids are often signatures of adaptive or reprogrammable systems in nature. Known examples of these systems, such as transcriptional activator-like effectors (TALE) and CRISPR, have been harnessed as powerful molecular tools with a wide range of applications including genome editing. The continued expansion of genomic sequence databases raises the possibility of prospectively identifying new such systems by computational mining. By leveraging sequence repeats as an organizing principle, here we develop a systematic genome mining approach to explore new types of naturally adaptive systems, five of which are discussed in greater detail. These results highlight the existence of a diverse range of intriguing systems in nature that remain to be explored and also provide a framework for future discovery efforts.

Information

Type
Research Article
Creative Commons
Creative Common License - CCCreative Common License - BYCreative Common License - NCCreative Common License - ND
This is an Open Access article, distributed under the terms of the Creative Commons Attribution-NonCommercial-NoDerivatives licence (http://creativecommons.org/licenses/by-nc-nd/4.0/), which permits non-commercial re-use, distribution, and reproduction in any medium, provided the original work is unaltered and is properly cited. The written permission of Cambridge University Press must be obtained for commercial re-use or in order to create a derivative work.
Copyright
© The Author(s) 2020. Published by Cambridge University Press
Figure 0

Fig. 1. Repeat structures in proteins and systems. (a) A comparison of different nucleic acid binding modules according to their modularity. Zinc Fingers, TALEs, and CRISPRs use repeats (dark grey), while TALEs and CRISPRs have hypervariable regions within their repeats that precisely determine the DNA binding specificity (red). (b) A schematic of different types of repeats and their diversification. (c) Basic mechanisms of diversification in prokaryotic genomes.

Figure 1

Fig. 2. Computational pipeline design for repeat protein analysis. (a) Schematic of protein-scale repeat pipeline. In the right most panel, N is the number of proteins containing a specific neighbouring repeat pair, while P (U,V) is the estimated joint distribution of neighbouring repeat pairs (u,v) obtained by counting the number of proteins with each specific pair of repeats and normalizing by the sum of all counts. The repeat rearrangement score, S, is the variation of information metric obtained by subtracting the mutual information, I (U,V), from the joint entropy, H (U,V). (b) Schematic of genome-scale repeat pipeline. The hypervariation score consists of computing an adjusted, non-redundant distance matrix between the hypervariable regions, and similarly for the constant regions. The hypervariation score, S, is the maximum ratio of the sum of the adjusted distance matrices over all hypervariable regions in the alignment. (c) Histogram of non-zero repeat rearrangement scores for hits from the protein-scale repeat pipeline, with an indicator for the score of the highest scoring TALE cluster. (d) Distribution of cluster sizes from the genome-scale repeat pipeline. (e) Distribution of the hypervariation score of all hits, with an indicator for score of the highest scoring TALE cluster. (f) Scatter plot of all within cluster percent identities and corresponding hypervariation score.

Figure 2

Fig. 3. Extensive signatures of modularity and recombination in a leucine-rich repeat (LRR) protein locus from Flavobacterium psychrophilum. (a) Graphical annotation of the LRR protein loci from ten strains of F. psychrophilum. (b) Domain architecture and sequence identity of a prototypical LRR protein (JIP02/86 #8). (c) Amino acid sequence logo of individual repeat units (n = 628) within intact LRR proteins. (d) Histogram of the number of repeat units within intact LRR proteins. (e) Structural model (trRosetta) of a prototypical LRR protein, highlighting the hypervariable positions (red) within the repeat units. The model was constructed from the first LRR protein in strain JIP02/86 (WP_011962357.1). (f) Size distribution of intact LRR proteins (red; n = 111) and protein fragments (blue; n = 206). (g) DNA microhomologies at high-confidence fragment-fragment junctions (left). Simulated microhomologies (right) based on random fragmentation of three intact loci (JIP02/86, CSF259-93, and FPG3).

Figure 3

Fig. 4. Leucine-rich repeat (LRR) proteins from Dictyostelium purpureum. (a) Repeat architectures of four representative D. purpureum LRR proteins. (b) Sequence logo of the LRR motifs. (c) Structural model of a representative LRR protein, with hypervariable residues shown as sticks. (d) Distribution of all pairs of hypervariable residues within a single LRR unit.

Figure 4

Fig. 5. (a) Splicing isoforms for the Solanum lycopersicum transcription factor LOC101240705 (Solyc02g091030). A majority of isoforms differ only in the displayed region containing a tandem array of amino acid repeats. (b) Top: sequence logo of the 12 amino acid repeats without deletions. Bottom: PSIPRED secondary structure prediction of a representative repeat.

Figure 5

Fig. 6. An array of serine proteases containing a hypervariable insert within the protease domain. (a) Graphical annotation of the protease locus from eight representative Streptosporangiceae strains. The hypervariable insert is shown in red. (b) Domain architecture and sequence identity of a prototypical protease (M. glauca #14). (c) Sequence logo of the catalytic serine and neighbouring residues from n = 223 proteases. (d) Histogram of hypervariable insert lengths. (e) Amino acid sequences of the inserts within the proteases from a representative locus (Herbidospora cretacea NBRC 15474). (f) Structural models (trRosetta) of representative proteases, constructed (left to right) from WP_061297158.1, WP_061297163.1, and WP_068929153.1.

Figure 6

Fig. 7. An array of alternating protein pairs from Photorhabdus containing localized variation. (a) Graphical annotation of seven representative loci from Photorhabdus species. (b) Domain architecture and sequence identity of a prototypical L protein (P. thracensis #2). (c) Amino acid sequences of the hypervariable inserts within the fifteen L proteins shown in (a). (de) Yeast two-hybrid assay for P. thracensis L–S protein interactions (HIS3 reporter). The L protein #2 from P. thracensis (WP_046976484.1) was used as the fixed scaffold for all L proteins in the assay. (f) Comparison with the ids gene cluster from Proteus mirabilis, which confers self-identity and social recognition (Gibbs et al., 2008). The genes idsB, idsC and idsF are shared between the Photorhabdus and P. mirabilis loci.

Supplementary material: File

Altae-Tran et al. supplementary material

Altae-Tran et al. supplementary material

Download Altae-Tran et al. supplementary material(File)
File 2.7 MB

Review: Computational Identification of Repeat-containing Proteins and Systems — R0/PR1

Conflict of interest statement

none.

Comments

Comments to Author: This paper describes the use of bioinformatics for identification of repeat containing proteins in nature. The authors developed a computational algorithm for identification of repetitive nucleotide sequence elements in proteins and nucleic acids. Such repetitive sequences are known in CRISPR and other systems, and they generally confer important functions. However, in many cases these functions are unknown or not even studied. The study presented here is interesting as it enables identification of repetitive sequences not earlier identified and it further discuss the possible function of four of these sequences. The paper is very well written and even though there is no experimental validation of any of the findings, the paper stands as a very strong novel contribution to the field. I have no specific comments to the paper, which can be accepted as is.

Review: Computational Identification of Repeat-containing Proteins and Systems — R0/PR2

Conflict of interest statement

Reviewer declares limited scientific exchange (unrelated to the current manuscript) with the corresponding author FZ, who is one of few world-leading experts in CRISPR-Cas.

Comments

Comments to Author: The manuscript "Computational Identification of Repeat-containing Proteins and Systems" describes a novel approach to mine entire genomic databases for repeating patterns of non-conserved amino acids in otherwise conserved proteins. This is an important insight. As nature has a preference of repurposing existing things, non-conserved areas in duplicated mostly homologous proteins could indicate an accelerated response to selective pressure, and thus lead to the discovery of new protein systems or functions. The authors present several examples of such hypervariability within various organisms, possibly constituting yet unknown modular interfaces adaptable to different targets. Analogous to antibodies, TALEs, or CRISPR, each discovery of a modular system means it can be reprogrammed artificially, which could have important and unforeseen applications in biotechnology and therapeutics.

The quality of the manuscript is excellent, and with increasing amount of available computing power, as well as more accessible DNA sequencing data, the manuscript could have a large scientific impact. This reviewer therefore recommends publication of the manuscript.

A few simple points:

Figure 2, the upper right formula seems incomplete.

Figure 4D could be slightly taller to avoid letters colliding.

Could the same approach be applied to DNA/RNA sequences themselves, for example, to find "RNA enzymes"?

Decision: Computational Identification of Repeat-containing Proteins and Systems — R0/PR3

Comments

Comments to Author: Reviewer #1: This paper describes the use of bioinformatics for identification of repeat containing proteins in nature. The authors developed a computational algorithm for identification of repetitive nucleotide sequence elements in proteins and nucleic acids. Such repetitive sequences are known in CRISPR and other systems, and they generally confer important functions. However, in many cases these functions are unknown or not even studied. The study presented here is interesting as it enables identification of repetitive sequences not earlier identified and it further discuss the possible function of four of these sequences. The paper is very well written and even though there is no experimental validation of any of the findings, the paper stands as a very strong novel contribution to the field. I have no specific comments to the paper, which can be accepted as is.

Reviewer #3: The manuscript "Computational Identification of Repeat-containing Proteins and Systems" describes a novel approach to mine entire genomic databases for repeating patterns of non-conserved amino acids in otherwise conserved proteins. This is an important insight. As nature has a preference of repurposing existing things, non-conserved areas in duplicated mostly homologous proteins could indicate an accelerated response to selective pressure, and thus lead to the discovery of new protein systems or functions. The authors present several examples of such hypervariability within various organisms, possibly constituting yet unknown modular interfaces adaptable to different targets. Analogous to antibodies, TALEs, or CRISPR, each discovery of a modular system means it can be reprogrammed artificially, which could have important and unforeseen applications in biotechnology and therapeutics.

The quality of the manuscript is excellent, and with increasing amount of available computing power, as well as more accessible DNA sequencing data, the manuscript could have a large scientific impact. This reviewer therefore recommends publication of the manuscript.

A few simple points:

Figure 2, the upper right formula seems incomplete.

Figure 4D could be slightly taller to avoid letters colliding.

Could the same approach be applied to DNA/RNA sequences themselves, for example, to find "RNA enzymes"?

Decision: Computational Identification of Repeat-containing Proteins and Systems — R1/PR4

Comments

No accompanying comment.