Options
Challenges of Automatically Detecting Offensive Language Online
Subtitle
Participation Paper for the Germeval Shared Task 2018 (HaUA)
Abstract
This paper presents our submission (HaUA) for Germeval Shared Task 1 (Binary Classification) on the identification of offensive language. With feature selection and features such as character ngrams, offensive word lexicons, and sentiment polarity, our SVM classifier is able to distinguish between offensive and nonoffensive German language tweets with an indomain F 1 score of 88.9%. In this paper, we report our methodology and discuss machine learning problems such as imbalance, overfitting, and the interpretability of machine learning algorithms. In the discussion section, we also briefly go beyond the technical perspectives and argue for a thorough discussion of the dilemma between internet security and freedom of speech, and what kind of language we are actually predicting with such algorithms.
Publication Type
ConferencePaper
Author •
De Smedt, Tom
Editor • •
Ruppendorfer, Josef
Siegel, Melanie
Wiegand, Michael
Date Issued
2018
Faculty
Institute / Institution
Published in
Proceedings of GermEval 2018 (KONVENS 2018)
Conference
14th Conference on Natural Language Processing (KONVENS 2018), Wien, 21.09.2018
Publisher
Verlag der Österreichischen Akademie der Wissenschaften
Publisher Place
Wien
Page Start
27
Page End
32
ISBN
978-3-7001-8435-5
Link to the original publication
HilPub short link