Unknown

Dataset Information

0

A curated mammography data set for use in computer-aided detection and diagnosis research.


ABSTRACT: Published research results are difficult to replicate due to the lack of a standard evaluation data set in the area of decision support systems in mammography; most computer-aided diagnosis (CADx) and detection (CADe) algorithms for breast cancer in mammography are evaluated on private data sets or on unspecified subsets of public databases. This causes an inability to directly compare the performance of methods or to replicate prior results. We seek to resolve this substantial challenge by releasing an updated and standardized version of the Digital Database for Screening Mammography (DDSM) for evaluation of future CADx and CADe systems (sometimes referred to generally as CAD) research in mammography. Our data set, the CBIS-DDSM (Curated Breast Imaging Subset of DDSM), includes decompressed images, data selection and curation by trained mammographers, updated mass segmentation and bounding boxes, and pathologic diagnosis for training data, formatted similarly to modern computer vision data sets. The data set contains 753 calcification cases and 891 mass cases, providing a data-set size capable of analyzing decision support systems in mammography.

SUBMITTER: Lee RS 

PROVIDER: S-EPMC5735920 | biostudies-literature | 2017 Dec

REPOSITORIES: biostudies-literature

altmetric image

Publications

A curated mammography data set for use in computer-aided detection and diagnosis research.

Lee Rebecca Sawyer RS   Gimenez Francisco F   Hoogi Assaf A   Miyake Kanae Kawai KK   Gorovoy Mia M   Rubin Daniel L DL  

Scientific data 20171219


Published research results are difficult to replicate due to the lack of a standard evaluation data set in the area of decision support systems in mammography; most computer-aided diagnosis (CADx) and detection (CADe) algorithms for breast cancer in mammography are evaluated on private data sets or on unspecified subsets of public databases. This causes an inability to directly compare the performance of methods or to replicate prior results. We seek to resolve this substantial challenge by rele  ...[more]

Similar Datasets

| S-EPMC3182841 | biostudies-literature
| S-EPMC4836172 | biostudies-literature
| S-EPMC3069993 | biostudies-other
| S-EPMC10182079 | biostudies-literature
| S-EPMC7987253 | biostudies-literature
| S-EPMC5557046 | biostudies-other
| S-EPMC7091549 | biostudies-literature
| S-EPMC8719464 | biostudies-literature
| S-EPMC8547575 | biostudies-literature
| S-EPMC10558498 | biostudies-literature