GPT Annotation Approach in the Application of Named Entity Recognition on Scopus Abstract Using SciBERT
Abstract
The rapid advancement of scientific research is reflected in the increasing number of publications on Scopus, which is the largest abstract database with a rigorous curation process. The abundance of abstracts highlights the need for automated methods to extract key information from scientific documents. This study uses the keyword “statistical analysis” in Scopus indexed publications for research with Indonesian institution affiliation and employs Named Entity Recognition to extract entities such as data, method, and task in order to understand the characteristics of the research. In the annotation process, we used GPT model as an annotator to generate silver data to train the SciBERT model. The GPT-4o model demonstrates more optimal performance compared to GPT-5 in terms of time and cost efficiency, with a Cohen’s kappa value of 0.75 for prompt version 3 using 20 labeled document examples. The SciBERT model trained using silver data achieves a f1-score of 52.94% with a precision value of 51.17% and a recall of 54.84%. Although the results show promising potential, the model still faces limitations in accurately detecting the boundaries and types of entities. Overall, using GPT to generate silver data can serve as an alternative in the data annotation process, particularly when resources are limited.