NewsCom-TOX: A corpus of comments on news articles annotated for toxicity in Spanish

Data de publicació

2025-04-02T16:11:33Z

2025-04-02T16:11:33Z

2024-01-17

2025-04-02T16:11:33Z

Resum

In this article, we present the NewsCom-TOX corpus, a new corpus manually annotated for toxicity in Spanish. NewsCom-TOX consists of 4359 comments in Spanish posted in response to 21 news articles on social media related to immigration, in order to analyse and identify messages with racial and xenophobic content. This corpus is multi-level annotated with different binary linguistic categories -stance, target, stereotype, sarcasm, mockery, insult, improper language, aggressiveness and intolerance- taking into account not only the information conveyed in each comment, but also the whole discourse thread in which the comment occurs, as well as the information conveyed in the news article, including their images. These categories allow us to identify the presence of toxicity and its intensity, that is, the level of toxicity of each comment. All this information is available for research purposes upon request. Here we describe the NewsCom-TOX corpus, the annotation tagset used, the criteria applied and the annotation process carried out, including the inter-annotator agreement tests conducted. A quantitative analysis of the results obtained is also provided. NewsCom-TOX is a linguistic resource that will be valuable for both linguistic and computational research in Spanish in NLP tasks for the detection of toxic information.

Tipus de document

Article


Versió acceptada

Llengua

Anglès

Publicat per

Springer Verlag

Documents relacionats

Versió postprint del document publicat a: https://doi.org/10.1007/s10579-023-09711-x

Language Resources And Evaluation, 2023, num.58, p. 1115-1155

https://doi.org/10.1007/s10579-023-09711-x

Citació recomanada

Aquesta citació s'ha generat automàticament.

Drets

(c) Springer Verlag, 2023

Aquest element apareix en la col·lecció o col·leccions següent(s)