How toxic is antisemitism? Potentials and limitations of automated toxicity scoring for antisemitic online content
作者: Helena Mihaljević, Elisabeth Steffen
分类: cs.CL, cs.AI, cs.CY
发布日期: 2023-10-05
备注: In: Proceedings of the 2nd Workshop on Computational Linguistics for Political Text Analysis (CPSS-2022), Potsdam, Germany, Sep 12, 2022
💡 一句话要点
评估在线反犹太主义内容的毒性评分方法及其局限性
🎯 匹配领域: 支柱一:机器人控制 (Robot Control)
关键词: 文本毒性评估 反犹太主义 内容审核 机器学习 社交媒体研究
📋 核心要点
- 现有的文本毒性评估方法在识别隐性反犹太主义内容时存在显著不足,导致内容审核效果不佳。
- 论文通过手动标注数据集,分析Perspective API在不同反犹太主义形式下的评分表现,揭示其局限性。
- 实验结果表明,API在识别显性反犹太主义时表现良好,但对隐性内容的评分准确性较低,且易被操控。
📝 摘要(中文)
本文探讨了Google和Jigsaw的Perspective API在检测反犹太主义在线内容中的潜力与局限性。通过对约3600条来自Telegram和Twitter的德语帖子进行手动标注,研究发现该API能够在基本层面上识别反犹太主义内容为有毒,但在处理隐性反犹太主义和持批判态度的文本时存在显著弱点。此外,简单的文本操作可以显著降低API评分,使得绕过基于该服务结果的内容审核变得相对容易。
🔬 方法详解
问题定义:本文旨在解决现有文本毒性评估方法在识别反犹太主义内容时的不足,尤其是隐性反犹太主义的识别困难和内容审核的有效性问题。
核心思路:通过构建一个手动标注的德语数据集,分析Perspective API对不同形式反犹太主义文本的评分,揭示其在内容审核中的潜在缺陷。
技术框架:研究采用了手动标注的语料库,包含来自Telegram和Twitter的帖子,利用Perspective API进行毒性评分,并进行对比分析。
关键创新:论文的创新点在于揭示了API在处理隐性反犹太主义内容时的弱点,并展示了如何通过文本操控降低评分,这在现有研究中尚未被充分探讨。
关键设计:研究中使用了简单的文本操作来测试API的鲁棒性,分析了不同反犹太主义形式的评分差异,重点关注API在处理批判性文本时的表现。
🖼️ 关键图片
📊 实验亮点
实验结果显示,Perspective API在识别显性反犹太主义内容时的准确率较高,但在处理隐性内容时评分显著降低,且通过简单文本操控可以使API评分下降,表明其在内容审核中的脆弱性。
🎯 应用场景
该研究的潜在应用场景包括社交媒体内容审核、在线社区管理以及反歧视政策的制定。通过改进毒性评分方法,可以更有效地识别和管理有害内容,促进网络环境的健康发展。
📄 摘要(原文)
The Perspective API, a popular text toxicity assessment service by Google and Jigsaw, has found wide adoption in several application areas, notably content moderation, monitoring, and social media research. We examine its potentials and limitations for the detection of antisemitic online content that, by definition, falls under the toxicity umbrella term. Using a manually annotated German-language dataset comprising around 3,600 posts from Telegram and Twitter, we explore as how toxic antisemitic texts are rated and how the toxicity scores differ regarding different subforms of antisemitism and the stance expressed in the texts. We show that, on a basic level, Perspective API recognizes antisemitic content as toxic, but shows critical weaknesses with respect to non-explicit forms of antisemitism and texts taking a critical stance towards it. Furthermore, using simple text manipulations, we demonstrate that the use of widespread antisemitic codes can substantially reduce API scores, making it rather easy to bypass content moderation based on the service's results.