Aug 2026· Journal of Computing and Information Science in Engineering· pp. 1-36· 0 citations
TL;DR
This study compares GPT-4's performance with that of 30 human participants from diverse backgrounds and shows that the CS/AI group had the most closely aligned responses with GPT-4.
Abstract
With large language models (LLMs) becoming ever more popular and their usage expanding into various domains, this study explores how effective an LLM like GPT-4 would be in analyzing requirement specification documents and using relevant information from those to create a tradespace matrix. This study compares GPT-4's performance with that of 30 human participants from diverse backgrounds, categorized into three groups: engineering, non-technical and (CS/AI). Each participant and the LLM completed a survey based on materials provided for evaluating a complex system design and populating a trade-space matrix. The analysis of these responses included within-group heatmaps, across-group comparisons, and human-versus-GPT-4 heatmap evaluations. The results show that the CS/AI group had the most closely aligned responses with GPT-4. This study demonstrates how LLMs can augment trade-space exploration and serve as an alternative in cases where employing a team of human experts from multiple backgrounds is not feasible.
No overall preference is indicated between human and machine-generated summaries; however, experts with greater familiarity with the Bulletin and higher educational attainment showed a marked preference for the official summaries.
Claudia Biancotti, C. Camassa, Marco Fruzzetti et al.· 0 citations
This work systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026, using staged screening and automated full-text coding to examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms.
GenAI can function as a research tool, but not as a substitute for methodological expertise, and has potential to increase efficiency of tasks which take advantage of its search and summarization abilities, as well as basic code debugging and algorithm formation.
Natalie Morosin, A. A. Nadi, Michael P. Wallace· 0 citations
A pilot project in which students in a statistics course within a data science engineering program created culturally diverse multiple-choice questions, generated answers using LLMs, and applied statistical methods to assess model accuracy is presented, supporting a cultural injection hypothesis.
Denis Iorga, Razvan Muntean, Mihai Masala et al.· Electronics· 0 citations
A Generative Pre-trained Transformers-only model that relies on prompt engineering against a Retrieval-Augmented Generation model that incorporates external university documents, specifically program flyers and a module handbook, integrated using Langchain are evaluated.
Meltem Cakar· Athens Journal of Τechnology...· 1 citation
Overall, the findings suggest that LLMs can approximate human coding in this case-specific setting, particularly at the coding level, and may serve as a scalable support tool for inductive qualitative analysis.
Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili et al.· Research Square· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.