Kernel methods and string kernels for authorship analysis

Popescu, Marius; Grozea, Cristian

2012

Conference Paper

Abstract

This paper presents our approach to the PAN 2012 Traditional Authorship Attribution tasks and the Sexual Predator Identification task. We approached these tasks with machine learning methods that work at the character level. More precisely, we treated texts as just sequences of symbols (strings) and used string kernels in conjunction with different kernel-based learning methods: supervised and unsupervised. The results were extremely good, we ranked first in most problem and overall in the traditional authorship attribution task, according to the evaluation provided by the organizers.

Author(s)

Popescu, Marius

Grozea, Cristian

Fraunhofer-Institut für Offene Kommunikationssysteme FOKUS

Mainwork

CLEF 2012 Conference and Labs of the Evaluation Forum. Evaluation Labs and Workshop. Online Working Notes

Conference

International Conference of the Cross-Language Evaluation Forum (CLEF) 2012

Options

Kernel methods and string kernels for authorship analysis