
Publication details
Publisher: Springer
Place: Berlin
Year: 2017
Pages: 274-286
Series: Lecture Notes in Computer Science
ISBN (Hardback): 9783319670072
Full citation:
, "A comparative study of language modeling to instance-based methods, and feature combinations for authorship attribution", in: Research and advanced technology for digital libraries, Berlin, Springer, 2017


A comparative study of language modeling to instance-based methods, and feature combinations for authorship attribution
pp. 274-286
in: Jaap Kamps, Giannis Tsakonas, Yannis Manolopoulos, Lazaros Iliadis, Ioannis Karydis (eds), Research and advanced technology for digital libraries, Berlin, Springer, 2017Abstract
We present a comparative study of language modeling to traditional instance-based methods for authorship attribution, using several different basic units as features, such as characters, words, and other simple lexical measurements, as well as we propose the use of part-of-speech (POS) tags as features for language modeling. In contrast to many other studies which focus on small sets of documents written by major writers regarding several topics, we consider a relatively large corpus with documents edited by non-professional writers regarding the same topic. We find that language models based on either characters or POS tags are the most effective, while the latter provide additional efficiency benefits and robustness against data sparsity. Moreover, we experiment with linearly combining several language models, as well as employing unions of several different feature types in instance-based methods. We find that both such combinations constitute viable strategies which generally improve effectiveness. By linearly combining three language models, based respectively on character, word, and POS trigrams, we achieve the best generalization accuracy of 96%.
Publication details
Publisher: Springer
Place: Berlin
Year: 2017
Pages: 274-286
Series: Lecture Notes in Computer Science
ISBN (Hardback): 9783319670072
Full citation:
, "A comparative study of language modeling to instance-based methods, and feature combinations for authorship attribution", in: Research and advanced technology for digital libraries, Berlin, Springer, 2017