Unknown

Dataset Information

0

Comprehensive analysis of the base composition around the transcription start site in Metazoa.


ABSTRACT:

Background

The transcription start site of a metazoan gene remains poorly understood, mostly because there is no clear signal present in all genes. Now that several sequenced metazoan genomes have been annotated, we have been able to compare the base composition around the transcription start site for all annotated genes across multiple genomes.

Results

The most prominent feature in the base compositions is a significant local variation in G+C content over a large region around the transcription start site. The change is present in all animal phyla but the extent of variation is different between distinct classes of vertebrates, and the shape of the variation is completely different between vertebrates and arthropods. Furthermore, the height of the variation correlates with CpG frequencies in vertebrates but not in invertebrates and it also correlates with gene expression, especially in mammals. We also detect GC and AT skews in all clades (where %G is not equal to %C or %A is not equal to %T respectively) but these occur in a more confined region around the transcription start site and in the coding region.

Conclusions

The dramatic changes in nucleotide composition in humans are a consequence of CpG nucleotide frequencies and of gene expression, the changes in Fugu could point to primordial CpG islands, and the changes in the fly are of a totally different kind and unrelated to dinucleotide frequencies.

SUBMITTER: Aerts S 

PROVIDER: S-EPMC436054 | biostudies-literature | 2004 Jun

REPOSITORIES: biostudies-literature

altmetric image

Publications

Comprehensive analysis of the base composition around the transcription start site in Metazoa.

Aerts Stein S   Thijs Gert G   Dabrowski Michal M   Moreau Yves Y   De Moor Bart B  

BMC genomics 20040601 1


<h4>Background</h4>The transcription start site of a metazoan gene remains poorly understood, mostly because there is no clear signal present in all genes. Now that several sequenced metazoan genomes have been annotated, we have been able to compare the base composition around the transcription start site for all annotated genes across multiple genomes.<h4>Results</h4>The most prominent feature in the base compositions is a significant local variation in G+C content over a large region around th  ...[more]

Similar Datasets

| S-EPMC1448210 | biostudies-literature
| S-EPMC3708499 | biostudies-literature
| S-EPMC6393241 | biostudies-literature
| S-EPMC5741190 | biostudies-literature
| S-EPMC4360255 | biostudies-literature
| S-EPMC7730318 | biostudies-literature
| S-EPMC3160847 | biostudies-literature
| S-EPMC3377991 | biostudies-literature
2013-11-11 | GSE49459 | GEO
2021-08-23 | GSE158048 | GEO