Options
From term extraction to lemma selection for an electronic LSP-dictionary in the field of mathematics
Abstract
We work on term extraction for a corpus-based LSP-dictionary. Our field of study is the mathematical domain of graph theory. Our working hypothesis is that mathematics lends itself to a specific approach for term and information extraction with a lexicographical purpose. We compare different methods for term extraction: The first one combines pattern-based and statistical means implemented by Schäfer et al. (2015), the second one has been developed especially for mathematical texts using domain-specific definition patterns based on work in the tradition of Meyer (2001). Further comparisons are made with a list of term candidates which are not part of the general language lexicon used in a version of TreeTagger trained on news text (Schmid, 1994) and with the term extraction provided by Sketch Engine (Kilgarriff et al., 2014). We use manual annotation by three expert raters and inter-rater agreement with κ-statistics to compare and evaluate the approaches. Additionally, we qualitatively analyse the extracted results. For selecting the lemmas, we work with a German corpus of lecture notes, textbooks and papers.
Publication Type
ConferencePaper
Author
Editor • • • • •
Kosem, Iztok
Cukr, Michal
Jakubíček, Milos
Kallas, Jelena
Krek, Simon
Tiberius, Carole
Date Issued
2021
Faculty
Institute / Institution
Published in
Electronic lexicography in the 21st century: post-editing lexicography: Proceedings of the eLex 2021 conference
Conference
7th Electronic lexicography in the 21st century conference, online, 05.07.-07.07.2021
Publisher
Lexical Computing
Publisher Place
Brno
Page Start
572
Page End
587
HilPub short link