Options
2027
Conference Paper
Title
Too Simple or Too Complex? Using Linguistic Signatures for AI-Generated Text Detection
Abstract
We studied the effectiveness of machine learning models and linguistic features for differentiating between multiple large language models (LLMs), addressing both binary and multi-class tasks. We introduce 188 linguistic features associated with text complexity to provide further insight into the differences among LLMs. Using multiple feature ranking and selection techniques, we identified the most discriminative features for training and conducted an exploratory analysis. Our experiments demonstrate strong performance in the binary classification scenario, achieving an in-domain F1 score of 96.3%. In the multi-class setting, text difficulty features improved the F1 score by 3.4% (in-domain). Our analysis revealed differences between LLMs, highlighting specific stylistic patterns and variances, such as type-token ratio, lexical diversity, and readability scores. AI-generated texts were identified through reduced lexical richness and variety in expression. The results highlight the potential of leveraging linguistic complexity features to improve the detection and characterisation of AI-generated text. An evaluation on three unseen test sets identified ongoing challenges in generalisability.
Author(s)