Recent advances in large language models have dramatically improved the quality of automatically generated text, enabling systems to produce coherent and contextually appropriate content across a wide range of domains. While these technologies support many beneficial applications, they also introduce risks related to misinformation and automated content manipulation. Detecting machine-generated text has therefore become an important research challenge. This paper investigates the effectiveness of transformer-based classifiers for identifying synthetic text. Two widely used architectures, BERT and RoBERTa, are evaluated within a supervised learning framework designed to distinguish between human-written and machine-generated documents. Experimental results show that contextual representations learned by transformer models capture linguistic patterns associated with automated text generation.
Sansonetti, G., Micarelli, A. (2026). Can We Reliably Detect AI-Generated Text? Evidence from Real-World Data. In Communications in Computer and Information Science (pp.156-165). Cham : Springer Science and Business Media Deutschland GmbH [10.1007/978-3-032-30836-8_18].
Can We Reliably Detect AI-Generated Text? Evidence from Real-World Data
Sansonetti, Giuseppe
;Micarelli, Alessandro
2026-01-01
Abstract
Recent advances in large language models have dramatically improved the quality of automatically generated text, enabling systems to produce coherent and contextually appropriate content across a wide range of domains. While these technologies support many beneficial applications, they also introduce risks related to misinformation and automated content manipulation. Detecting machine-generated text has therefore become an important research challenge. This paper investigates the effectiveness of transformer-based classifiers for identifying synthetic text. Two widely used architectures, BERT and RoBERTa, are evaluated within a supervised learning framework designed to distinguish between human-written and machine-generated documents. Experimental results show that contextual representations learned by transformer models capture linguistic patterns associated with automated text generation.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


