Hybrid Learning for AI-Generated Text Detection and Attribution

No Thumbnail Available

Date

2026

Journal Title

Journal ISSN

Volume Title

Publisher

university of ghardaia

Abstract

The rapid advancement of large language models (LLMs) has significantly in- creased the quality and fluency of AI-generated text, making it increasingly difficult to distinguish from human-written content. This challenge is further amplified in source attribution tasks, where texts generated by multiple LLMs often exhibit overlapping linguistic and stylistic characteristics. This study proposes a hybrid framework for AI-generated text detection and attribution by combining semantic and stylometric representations to improve classification performance. The framework integrates con- textual semantic embeddings extracted using RoBERTa with handcrafted stylometric features, while supervised contrastive learning is employed to enhance feature discrim- inability and class separability. The proposed approach addresses both binary classi- fication (human versus AI-generated text) and multi-class attribution for identifying the generating model. The fused representation is processed through a multilayer per- ceptron (MLP) classifier for final prediction. Experimental evaluations conducted on the MAGE Dataset demonstrate the effectiveness of the proposed framework, achiev- ing a 93.20% macro F1-score in binary classification and an 82.22% macro F1-score in multi-class attribution. Comparative results show that the proposed model outper- forms traditional machine learning baselines, including Logistic Regression, Random Forest, Support Vector Machine (SVM), demonstrating its robustness and effectiveness for AI-generated text detection and attribution tasks.

Description

Specialty: Intelligent Systems for Knowledge Extraction Slimane Bellaouar/encadreur

Keywords

AI-generated text detection, large language models (LLMs), source attribution, contrastive learning.

Citation

Endorsement

Review

Supplemented By

Referenced By