Dataset Information

Privacy-preserving search for chemical compound databases.

ABSTRACT:

Background

Searching for similar compounds in a database is the most important process for in-silico drug screening. Since a query compound is an important starting point for the new drug, a query holder, who is afraid of the query being monitored by the database server, usually downloads all the records in the database and uses them in a closed network. However, a serious dilemma arises when the database holder also wants to output no information except for the search results, and such a dilemma prevents the use of many important data resources.

Results

In order to overcome this dilemma, we developed a novel cryptographic protocol that enables database searching while keeping both the query holder's privacy and database holder's privacy. Generally, the application of cryptographic techniques to practical problems is difficult because versatile techniques are computationally expensive while computationally inexpensive techniques can perform only trivial computation tasks. In this study, our protocol is successfully built only from an additive-homomorphic cryptosystem, which allows only addition performed on encrypted values but is computationally efficient compared with versatile techniques such as general purpose multi-party computation. In an experiment searching ChEMBL, which consists of more than 1,200,000 compounds, the proposed method was 36,900 times faster in CPU time and 12,000 times as efficient in communication size compared with general purpose multi-party computation.

Conclusion

We proposed a novel privacy-preserving protocol for searching chemical compound databases. The proposed method, easily scaling for large-scale databases, may help to accelerate drug discovery research by making full use of unused but valuable data that includes sensitive information.

SUBMITTER: Shimizu K

PROVIDER: S-EPMC4704467 | biostudies-literature | 2015

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Privacy-preserving search for chemical compound databases.

Shimizu Kana K Nuida Koji K Arai Hiromi H Mitsunari Shigeo S Attrapadung Nuttapong N Hamada Michiaki M Tsuda Koji K Hirokawa Takatsugu T Sakuma Jun J Hanaoka Goichiro G Asai Kiyoshi K

BMC bioinformatics 20151209

<h4>Background</h4>Searching for similar compounds in a database is the most important process for in-silico drug screening. Since a query compound is an important starting point for the new drug, a query holder, who is afraid of the query being monitored by the database server, usually downloads all the records in the database and uses them in a closed network. However, a serious dilemma arises when the database holder also wants to output no information except for the search results, and such ...[more]

PMID: 26678650

Dataset Information

Privacy-preserving search for chemical compound databases.

Background

Results

Conclusion

Publications

Privacy-preserving search for chemical compound databases.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets

Similar Datasets

Privacy-preserving record linkage in large databases using secure multiparty computation.
| S-EPMC6180364 | biostudies-other

Privacy-preserving <i>k</i>-NN interpolation over two encrypted databases.
| S-EPMC9202633 | biostudies-literature

Efficient privacy-preserving string search and an application in genomics.
| S-EPMC4892414 | biostudies-literature

ACE: A Consent-Embedded privacy-preserving search on genomic database.
| S-EPMC11636870 | biostudies-literature

Splitting chemical structure data sets for federated privacy-preserving machine learning.
| S-EPMC8650276 | biostudies-literature

Privacy-Preserving Process Mining in Healthcare.
| S-EPMC7084661 | biostudies-literature

Privacy preserving interactive record linkage (PPIRL).
| S-EPMC3932473 | biostudies-literature

Statistical-based database fingerprint: chemical space dependent representation of compound databases.
| S-EPMC6755589 | biostudies-literature

Virtual Screening of Natural Chemical Databases to Search for Potential ACE2 Inhibitors.
| S-EPMC8911956 | biostudies-literature

Identifying compound-target associations by combining bioactivity profile similarity search and public databases mining.
| S-EPMC3180241 | biostudies-literature