Options
Automatic Content-based Categorization of Wikipedia Articles
Abstract
Wikipedia’s article contents and its category hierarchy are widely used to produce semantic resources which improve performance on tasks like text classification and keyword extraction. The reverse – using text classification methods for predicting the categories of Wikipedia articles – has attracted less attention so far. We propose to “return the favor” and use text classifiers to improve Wikipedia. This could support the emergence of a virtuous circle between the wisdom of the crowds and machine learning/NLP methods.
We define the categorization of Wikipedia articles as a multi-label classification task, describe two solutions to the task, and perform experiments that show that our approach is feasible despite the high number of labels.
We define the categorization of Wikipedia articles as a multi-label classification task, describe two solutions to the task, and perform experiments that show that our approach is feasible despite the high number of labels.
Publication Type
ConferencePaper
Author •
Gantner, Zeno
Date Issued
2009
Faculty
Institute / Institution
Published in
Proceedings of the 2009 Workshop on the People’s Web Meets NLP, ACL-IJCNLP 2009
Conference
2009 Workshop on the People’s Web Meets NLP, ACL-IJCNLP 2009, Singapur, 07.08.2009
Publisher
ACL
Publisher Place
Stroudsburg
Page Start
32
Page End
37
ISBN
978-1-932432-55-8
Link to the original publication
HilPub short link