Unknown

Dataset Information

0

Codon optimization with deep learning to enhance protein expression.


ABSTRACT: Heterologous expression is the main approach for recombinant protein production ingenetic synthesis, for which codon optimization is necessary. The existing optimization methods are based on biological indexes. In this paper, we propose a novel codon optimization method based on deep learning. First, we introduce the concept of codon boxes, via which DNA sequences can be recoded into codon box sequences while ignoring the order of bases. Then, the problem of codon optimization can be converted to sequence annotation of corresponding amino acids with codon boxes. The codon optimization models for Escherichia Coli were trained by the Bidirectional Long-Short-Term Memory Conditional Random Field. Theoretically, deep learning is a good method to obtain the distribution characteristics of DNA. In addition to the comparison of the codon adaptation index, protein expression experiments for plasmodium falciparum candidate vaccine and polymerase acidic protein were implemented for comparison with the original sequences and the optimized sequences from Genewiz and ThermoFisher. The results show that our method for enhancing protein expression is efficient and competitive.

SUBMITTER: Fu H 

PROVIDER: S-EPMC7572362 | biostudies-literature | 2020 Oct

REPOSITORIES: biostudies-literature

altmetric image

Publications

Codon optimization with deep learning to enhance protein expression.

Fu Hongguang H   Liang Yanbing Y   Zhong Xiuqin X   Pan ZhiLing Z   Huang Lei L   Zhang HaiLin H   Xu Yang Y   Zhou Wei W   Liu Zhong Z  

Scientific reports 20201019 1


Heterologous expression is the main approach for recombinant protein production ingenetic synthesis, for which codon optimization is necessary. The existing optimization methods are based on biological indexes. In this paper, we propose a novel codon optimization method based on deep learning. First, we introduce the concept of codon boxes, via which DNA sequences can be recoded into codon box sequences while ignoring the order of bases. Then, the problem of codon optimization can be converted t  ...[more]

Similar Datasets

| S-EPMC3495653 | biostudies-literature
| S-EPMC3056764 | biostudies-literature
| S-EPMC3048298 | biostudies-literature
2024-02-03 | GSE254493 | GEO
| S-EPMC4899721 | biostudies-literature
| S-EPMC7757209 | biostudies-literature
| S-EPMC11200022 | biostudies-literature
| S-EPMC6656766 | biostudies-literature
| S-EPMC6786543 | biostudies-literature
| S-EPMC9951436 | biostudies-literature