Unknown

Dataset Information

0

Measuring Global Disease with Wikipedia: Success, Failure, and a Research Agenda.


ABSTRACT: Effective disease monitoring provides a foundation for effective public health systems. This has historically been accomplished with patient contact and bureaucratic aggregation, which tends to be slow and expensive. Recent internet-based approaches promise to be real-time and cheap, with few parameters. However, the question of when and how these approaches work remains open. We addressed this question using Wikipedia access logs and category links. Our experiments, replicable and extensible using our open source code and data, test the effect of semantic article filtering, amount of training data, forecast horizon, and model staleness by comparing across 6 diseases and 4 countries using thousands of individual models. We found that our minimal-configuration, language-agnostic article selection process based on semantic relatedness is effective for improving predictions, and that our approach is relatively insensitive to the amount and age of training data. We also found, in contrast to prior work, very little forecasting value, and we argue that this is consistent with theoretical considerations about the nature of forecasting. These mixed results lead us to propose that the currently observational field of internet-based disease surveillance must pivot to include theoretical models of information flow as well as controlled experiments based on simulations of disease.

SUBMITTER: Priedhorsky R 

PROVIDER: S-EPMC5542563 | biostudies-literature | 2017 Feb-Mar

REPOSITORIES: biostudies-literature

altmetric image

Publications

Measuring Global Disease with Wikipedia: Success, Failure, and a Research Agenda.

Priedhorsky Reid R   Osthus Dave D   Daughton Ashlynn R AR   Moran Kelly R KR   Generous Nicholas N   Fairchild Geoffrey G   Deshpande Alina A   Del Valle Sara Y SY  

CSCW : proceedings of the Conference on Computer-Supported Cooperative Work. Conference on Computer-Supported Cooperative Work 20170201


Effective disease monitoring provides a foundation for effective public health systems. This has historically been accomplished with patient contact and bureaucratic aggregation, which tends to be slow and expensive. Recent internet-based approaches promise to be real-time and cheap, with few parameters. However, the question of <i>when and how</i> these approaches work remains open. We addressed this question using Wikipedia access logs and category links. Our experiments, replicable and extens  ...[more]

Similar Datasets

| S-EPMC6075892 | biostudies-other
| S-EPMC4231164 | biostudies-literature
| S-EPMC3144035 | biostudies-literature
| S-EPMC5409523 | biostudies-literature
| S-EPMC5798175 | biostudies-literature
| S-EPMC6429801 | biostudies-other
| S-EPMC2671479 | biostudies-literature
| S-EPMC8160487 | biostudies-literature
| S-EPMC7056414 | biostudies-literature
| S-EPMC6851449 | biostudies-literature