Hybrid Learning for AI-Generated Text Detection and Attribution
No Thumbnail Available
Date
2026
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
university of ghardaia
Abstract
The rapid advancement of large language models (LLMs) has significantly in-
creased the quality and fluency of AI-generated text, making it increasingly difficult to
distinguish from human-written content. This challenge is further amplified in source
attribution tasks, where texts generated by multiple LLMs often exhibit overlapping
linguistic and stylistic characteristics. This study proposes a hybrid framework for
AI-generated text detection and attribution by combining semantic and stylometric
representations to improve classification performance. The framework integrates con-
textual semantic embeddings extracted using RoBERTa with handcrafted stylometric
features, while supervised contrastive learning is employed to enhance feature discrim-
inability and class separability. The proposed approach addresses both binary classi-
fication (human versus AI-generated text) and multi-class attribution for identifying
the generating model. The fused representation is processed through a multilayer per-
ceptron (MLP) classifier for final prediction. Experimental evaluations conducted on
the MAGE Dataset demonstrate the effectiveness of the proposed framework, achiev-
ing a 93.20% macro F1-score in binary classification and an 82.22% macro F1-score
in multi-class attribution. Comparative results show that the proposed model outper-
forms traditional machine learning baselines, including Logistic Regression, Random
Forest, Support Vector Machine (SVM), demonstrating its robustness and effectiveness
for AI-generated text detection and attribution tasks.
Description
Specialty: Intelligent Systems for Knowledge Extraction
Slimane Bellaouar/encadreur
Keywords
AI-generated text detection, large language models (LLMs), source attribution, contrastive learning.
