Статья

When Automated Assessment Meets Automated Content Generation: Examining Text Quality in the Era of GPTs

Marialena BevilacquaUniversity of Notre Dame, Notre Dame, Indiana, USAKezia OketchUniversity of Notre Dame, Notre Dame, Indiana, USARuiyang QinUniversity of Notre Dame, Notre Dame, Indiana, USAWill StameyUniversity of Notre Dame, Notre Dame, Indiana, USAXinyuan ZhangUniversity of Notre Dame, Notre Dame, Indiana, USAYi GanGeorgia Tech, Atlanta, Georgia, USAKai YangShenzhen University, Shenzhen, ChinaAhmed AbbasiUniversity of Notre Dame, Notre Dame, Indiana, USA

2024en

ABI

Аннотация

The use of machine learning (ML) models to assess and score textual data has become increasingly pervasive in an array of contexts including natural language processing, information retrieval, search and recommendation, and credibility assessment of online content. A significant disruption at the intersection of ML and text are text-generating large-language models (LLMs) such as generative pre-trained transformers (GPTs). We empirically assess the differences in how ML-based scoring models trained on human content assess the quality of content generated by humans versus GPTs. To do so, we propose an analysis framework that encompasses essay scoring ML models, human- and ML-generated essays, and a statistical model that parsimoniously considers the impact of type of respondent, prompt genre, and the ML model used for assessment model. A rich testbed is utilized that encompasses 18,460 human-generated and GPT-based essays. Results of our benchmark analysis reveal that LLMs and transformer pretrained language models (PLMs) more accurately score human essay quality as compared to CNN/RNN and feature-based ML methods. Interestingly, we find that LLMs and transformer PLMs tend to score GPT-generated text 10–20% higher on average, relative to human-authored documents. Conversely, traditional deep learning and feature-based ML models score human text considerably higher. Further analysis reveals that even though the LLMs and transformer PLMs are exclusively fine-tuned on human text, they more prominently attend to certain tokens appearing only in GPT-generated text, possibly (in part) due to familiarity/overlap in pre-training. Our framework and results have implications for text classification settings where automated scoring of text is likely to be disrupted by generative AI.

Перевод пока недоступен

Идентификаторы

DOI: 10.1145/3702639

Цитирования и источники

Цитирований: 2Использованных источников: 0

Показатели — AkademScholar