Text Classification Using Large Language Models to Detect AI and Human-Generated Sentences
Authors
| Issue | Vol. 12 No. 1 (2026) |
| Published | 4 August 2026 |
| Section | Articles |
| Pages | 66-77 |
Abstract
With the rapid development of the technology of Artificial Intelligence (AI), especially Large Language Models (LLMs) like ChatGPT, GPT-4 and Gemini, systems have become able to generate texts very similar to human writing. This similarity has been a boon to various sectors, but also poses fresh challenges of content authenticity and data integrity. A significant challenge is finding a way to automatically and accurately distinguish AI-generated sentences from human-written ones. This research focuses on building a text classification model that uses Large Language Models to distinguish between AI-generated and human-written sentences. This approach of research is based on recent research models that combine deep learning and classification using LLMs . The research process involves gathering human and AI-generated text data, pre-processing the text to normalize and tokenize it, feature extraction using the embeddings of large language models like IndoBERT, training the binary classification model, and assessing the model's performance using the metrics accuracy, precision, recall, and F1-score.
Keywords: Text clustering, Large language models, AI text identification, Text authenticity.
