Hier finden Sie wissenschaftliche Publikationen aus den Fraunhofer-Instituten.

Kernel methods and string kernels for authorship analysis

Notebook for PAN at CLEF 2012
: Popescu, Marius; Grozea, Cristian

Volltext (PDF; )

Forner, P.:
CLEF 2012 Conference and Labs of the Evaluation Forum. Evaluation Labs and Workshop. Online Working Notes : Information Access Evaluation meets Multilinguality, Multimodality, and Visual Analytics, Rome, Italy, September 17-20, 2012
Rome, 2012 (CEUR Workshop Proceedings 1178)
ISBN: 978-88-904810-3-1
12 S.
International Conference of the Cross-Language Evaluation Forum (CLEF) <3, 2012, Rome>
Konferenzbeitrag, Elektronische Publikation
Fraunhofer FOKUS ()
authorship analysis; natural language processing; string kernels; kernel methods; machine learning

This paper presents our approach to the PAN 2012 Traditional Authorship Attribution tasks and the Sexual Predator Identification task. We approached these tasks with machine learning methods that work at the character level. More precisely, we treated texts as just sequences of symbols (strings) and used string kernels in conjunction with different kernel-based learning methods: supervised and unsupervised. The results were extremely good, we ranked first in most problem and overall in the traditional authorship attribution task, according to the evaluation provided by the organizers.