Unknown

Dataset Information

0

Searching for transcription factor binding sites in vector spaces.


ABSTRACT:

Background

Computational approaches to transcription factor binding site identification have been actively researched in the past decade. Learning from known binding sites, new binding sites of a transcription factor in unannotated sequences can be identified. A number of search methods have been introduced over the years. However, one can rarely find one single method that performs the best on all the transcription factors. Instead, to identify the best method for a particular transcription factor, one usually has to compare a handful of methods. Hence, it is highly desirable for a method to perform automatic optimization for individual transcription factors.

Results

We proposed to search for transcription factor binding sites in vector spaces. This framework allows us to identify the best method for each individual transcription factor. We further introduced two novel methods, the negative-to-positive vector (NPV) and optimal discriminating vector (ODV) methods, to construct query vectors to search for binding sites in vector spaces. Extensive cross-validation experiments showed that the proposed methods significantly outperformed the ungapped likelihood under positional background method, a state-of-the-art method, and the widely-used position-specific scoring matrix method. We further demonstrated that motif subtypes of a TF can be readily identified in this framework and two variants called the k NPV and k ODV methods benefited significantly from motif subtype identification. Finally, independent validation on ChIP-seq data showed that the ODV and NPV methods significantly outperformed the other compared methods.

Conclusions

We conclude that the proposed framework is highly flexible. It enables the two novel methods to automatically identify a TF-specific subspace to search for binding sites. Implementations are available as source code at: http://biogrid.engr.uconn.edu/tfbs_search/.

SUBMITTER: Lee C 

PROVIDER: S-EPMC3543194 | biostudies-literature | 2012 Aug

REPOSITORIES: biostudies-literature

altmetric image

Publications

Searching for transcription factor binding sites in vector spaces.

Lee Chih C   Huang Chun-Hsi CH  

BMC bioinformatics 20120827


<h4>Background</h4>Computational approaches to transcription factor binding site identification have been actively researched in the past decade. Learning from known binding sites, new binding sites of a transcription factor in unannotated sequences can be identified. A number of search methods have been introduced over the years. However, one can rarely find one single method that performs the best on all the transcription factors. Instead, to identify the best method for a particular transcrip  ...[more]

Similar Datasets

| S-EPMC2647310 | biostudies-literature
| S-EPMC1261160 | biostudies-literature
| S-EPMC3898213 | biostudies-literature
| S-EPMC2588498 | biostudies-literature
| S-EPMC2824720 | biostudies-literature
| S-EPMC4091707 | biostudies-literature
| S-EPMC3931712 | biostudies-literature
| S-EPMC9358416 | biostudies-literature
| S-EPMC3244768 | biostudies-literature
| S-EPMC3522170 | biostudies-literature