Dataset Information

Hate speech detection and racial bias mitigation in social media based on BERT model.

ABSTRACT: Disparate biases associated with datasets and trained classifiers in hateful and abusive content identification tasks have raised many concerns recently. Although the problem of biased datasets on abusive language detection has been addressed more frequently, biases arising from trained classifiers have not yet been a matter of concern. In this paper, we first introduce a transfer learning approach for hate speech detection based on an existing pre-trained language model called BERT (Bidirectional Encoder Representations from Transformers) and evaluate the proposed model on two publicly available datasets that have been annotated for racism, sexism, hate or offensive content on Twitter. Next, we introduce a bias alleviation mechanism to mitigate the effect of bias in training set during the fine-tuning of our pre-trained BERT-based model for hate speech detection. Toward that end, we use an existing regularization method to reweight input samples, thereby decreasing the effects of high correlated training set' s n-grams with class labels, and then fine-tune our pre-trained BERT-based model with the new re-weighted samples. To evaluate our bias alleviation mechanism, we employed a cross-domain approach in which we use the trained classifiers on the aforementioned datasets to predict the labels of two new datasets from Twitter, AAE-aligned and White-aligned groups, which indicate tweets written in African-American English (AAE) and Standard American English (SAE), respectively. The results show the existence of systematic racial bias in trained classifiers, as they tend to assign tweets written in AAE from AAE-aligned group to negative classes such as racism, sexism, hate, and offensive more often than tweets written in SAE from White-aligned group. However, the racial bias in our classifiers reduces significantly after our bias alleviation mechanism is incorporated. This work could institute the first step towards debiasing hate speech and abusive language detection systems.

SUBMITTER: Mozafari M

PROVIDER: S-EPMC7451563 | biostudies-literature | 2020

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Hate speech detection and racial bias mitigation in social media based on BERT model.

Mozafari Marzieh M Farahbakhsh Reza R Crespi Noël N

PloS one 20200827 8

Disparate biases associated with datasets and trained classifiers in hateful and abusive content identification tasks have raised many concerns recently. Although the problem of biased datasets on abusive language detection has been addressed more frequently, biases arising from trained classifiers have not yet been a matter of concern. In this paper, we first introduce a transfer learning approach for hate speech detection based on an existing pre-trained language model called BERT (Bidirection ...[more]

PMID: 32853205

Dataset Information

Hate speech detection and racial bias mitigation in social media based on BERT model.

Publications

Hate speech detection and racial bias mitigation in social media based on BERT model.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets

Similar Datasets

A curated dataset for hate speech detection on social media text.
| S-EPMC9807815 | biostudies-literature

Moralized language predicts hate speech on social media.
| S-EPMC9837664 | biostudies-literature

Postgraduate psychological stress detection from social media using BERT-Fused model.
| S-EPMC11527284 | biostudies-literature

Examining Exposure to Messaging, Content, and Hate Speech from Partisan News Social Media Posts on Racial and Ethnic Health Disparities.
| S-EPMC9960309 | biostudies-literature

Empathy-based counterspeech can reduce racist hate speech in a social media field experiment.
| S-EPMC8685915 | biostudies-literature

Hate speech detection: Challenges and solutions.
| S-EPMC6701757 | biostudies-literature

Stance detection with BERT embeddings for credibility analysis of information on social media.
| S-EPMC8053013 | biostudies-literature

Is hate speech detection the solution the world wants?
| S-EPMC10013846 | biostudies-literature

Hate speech detection in the Arabic language: corpus design, construction, and evaluation
| S-EPMC10912174 | biostudies-literature

The risk of racial bias while tracking influenza-related content on social media using machine learning.
| S-EPMC7973478 | biostudies-literature