Dataset Information

An Analysis of French-Language Tweets About COVID-19 Vaccines: Supervised Learning Approach.

ABSTRACT:

Background

As the COVID-19 pandemic progressed, disinformation, fake news, and conspiracy theories spread through many parts of society. However, the disinformation spreading through social media is, according to the literature, one of the causes of increased COVID-19 vaccine hesitancy. In this context, the analysis of social media posts is particularly important, but the large amount of data exchanged on social media platforms requires specific methods. This is why machine learning and natural language processing models are increasingly applied to social media data.

Objective

The aim of this study is to examine the capability of the CamemBERT French-language model to faithfully predict the elaborated categories, with the knowledge that tweets about vaccination are often ambiguous, sarcastic, or irrelevant to the studied topic.

Methods

A total of 901,908 unique French-language tweets related to vaccination published between July 12, 2021, and August 11, 2021, were extracted using Twitter's application programming interface (version 2; Twitter Inc). Approximately 2000 randomly selected tweets were labeled with 2 types of categorizations: (1) arguments for (pros) or against (cons) vaccination (health measures included) and (2) type of content (scientific, political, social, or vaccination status). The CamemBERT model was fine-tuned and tested for the classification of French-language tweets. The model's performance was assessed by computing the F1-score, and confusion matrices were obtained.

Results

The accuracy of the applied machine learning reached up to 70.6% for the first classification (pro and con tweets) and up to 90% for the second classification (scientific and political tweets). Furthermore, a tweet was 1.86 times more likely to be incorrectly classified by the model if it contained fewer than 170 characters (odds ratio 1.86; 95% CI 1.20-2.86).

Conclusions

The accuracy of the model is affected by the classification chosen and the topic of the message examined. When the vaccine debate is jostled by contested political decisions, tweet content becomes so heterogeneous that the accuracy of the model drops for less differentiated classes. However, our tests showed that it is possible to improve the accuracy by selecting tweets using a new method based on tweet length.

SUBMITTER: Sauvayre R

PROVIDER: S-EPMC9116457 | biostudies-literature |

REPOSITORIES: biostudies-literature

ACCESS DATA

Similar Datasets

Project description:BackgroundFamily violence (including intimate partner violence/domestic violence, child abuse, and elder abuse) is a hidden pandemic happening alongside COVID-19. The rates of family violence are rising fast, and women and children are disproportionately affected and vulnerable during this time.ObjectiveThis study aims to provide a large-scale analysis of public discourse on family violence and the COVID-19 pandemic on Twitter.MethodsWe analyzed over 1 million tweets related to family violence and COVID-19 from April 12 to July 16, 2020. We used the machine learning approach Latent Dirichlet Allocation and identified salient themes, topics, and representative tweets.ResultsWe extracted 9 themes from 1,015,874 tweets on family violence and the COVID-19 pandemic: (1) increased vulnerability: COVID-19 and family violence (eg, rising rates, increases in hotline calls, homicide); (2) types of family violence (eg, child abuse, domestic violence, sexual abuse); (3) forms of family violence (eg, physical aggression, coercive control); (4) risk factors linked to family violence (eg, alcohol abuse, financial constraints, guns, quarantine); (5) victims of family violence (eg, the LGBTQ [lesbian, gay, bisexual, transgender, and queer or questioning] community, women, women of color, children); (6) social services for family violence (eg, hotlines, social workers, confidential services, shelters, funding); (7) law enforcement response (eg, 911 calls, police arrest, protective orders, abuse reports); (8) social movements and awareness (eg, support victims, raise awareness); and (9) domestic violence-related news (eg, Tara Reade, Melissa DeRosa).ConclusionsThis study overcomes limitations in the existing scholarship where data on the consequences of COVID-19 on family violence are lacking. We contribute to understanding family violence during the pandemic by providing surveillance via tweets. This is essential for identifying potentially useful policy programs that can offer targeted support for victims and survivors as we prepare for future outbreaks.

Project description:Social media is increasingly being used to express opinions and attitudes toward vaccines. The vaccine stance of social media posts can be classified in almost real-time using machine learning. We describe the use of a Transformer-based machine learning model for analyzing vaccine stance of Italian tweets, and demonstrate the need to address changes over time in vaccine-related language, through periodic model retraining. Vaccine-related tweets were collected through a platform developed for the European Joint Action on Vaccination. Two datasets were collected, the first between November 2019 and June 2020, the second from April to September 2021. The tweets were manually categorized by three independent annotators. After cleaning, the total dataset consisted of 1,736 tweets with 3 categories (promotional, neutral, and discouraging). The manually classified tweets were used to train and test various machine learning models. The model that classified the data most similarly to humans was XLM-Roberta-large, a multilingual version of the Transformer-based model RoBERTa. The model hyper-parameters were tuned and then the model ran five times. The fine-tuned model with the best F-score over the validation dataset was selected. Running the selected fine-tuned model on just the first test dataset resulted in an accuracy of 72.8% (F-score 0.713). Using this model on the second test dataset resulted in a 10% drop in accuracy to 62.1% (F-score 0.617), indicating that the model recognized a difference in language between the datasets. On the combined test datasets the accuracy was 70.1% (F-score 0.689). Retraining the model using data from the first and second datasets increased the accuracy over the second test dataset to 71.3% (F-score 0.713), a 9% improvement from when using just the first dataset for training. The accuracy over the first test dataset remained the same at 72.8% (F-score 0.721). The accuracy over the combined test datasets was then 72.4% (F-score 0.720), a 2% improvement. Through fine-tuning a machine-learning model on task-specific data, the accuracy achieved in categorizing tweets was close to that expected by a single human annotator. Regular training of machine-learning models with recent data is advisable to maximize accuracy.