Unknown

Dataset Information

0

Protein topology and stability define the space of allowed sequences.


ABSTRACT: We describe a new approach to explore and quantify the sequence space associated with a given protein structure. A set of sequences are optimized for a given target structure, using all-atom models and a physical energy function. Specificity of the sequence for its target is ensured by using the random energy model, which keeps the amino acid composition of the sequence constant. The designed sequences provide a multiple sequence alignment that describes the sequence space compatible with the structure of interest; here the size of this space is estimated by using an information entropy measure. In parallel, multiple alignments of naturally occurring sequences can be derived by using either sequence or structure alignments. We compared these 3 independent multiple sequence alignments for 10 different proteins, ranging in size from 56 to 310 residues. We observed that the subset of the sequence space derived by using our design procedure is similar in size to the sequence spaces observed in nature. These results suggest that the volume of sequence space compatible with a given protein fold is defined by the length of the protein as well as by the topology (i.e., geometry of the polypeptide chain) and the stability (i.e., free energy of denaturation) of the fold.

SUBMITTER: Koehl P 

PROVIDER: S-EPMC122181 | biostudies-literature | 2002 Feb

REPOSITORIES: biostudies-literature

altmetric image

Publications

Protein topology and stability define the space of allowed sequences.

Koehl Patrice P   Levitt Michael M  

Proceedings of the National Academy of Sciences of the United States of America 20020122 3


We describe a new approach to explore and quantify the sequence space associated with a given protein structure. A set of sequences are optimized for a given target structure, using all-atom models and a physical energy function. Specificity of the sequence for its target is ensured by using the random energy model, which keeps the amino acid composition of the sequence constant. The designed sequences provide a multiple sequence alignment that describes the sequence space compatible with the st  ...[more]

Similar Datasets

| S-EPMC7505158 | biostudies-literature
| S-EPMC11199558 | biostudies-literature
| S-EPMC3851920 | biostudies-literature
| S-EPMC10912673 | biostudies-literature
| S-EPMC2474553 | biostudies-literature
| S-EPMC8880924 | biostudies-literature
| S-EPMC4569382 | biostudies-literature
| S-EPMC3277656 | biostudies-literature
| S-EPMC7794440 | biostudies-literature
| S-EPMC5958423 | biostudies-literature