Conversational generative models are increasingly used to produce, transform, and revise information through multi-turn exchanges. However, large human–AI datasets are still commonly analyzed through performance, preference, or content lenses, leaving conversation-level behavioral structure less operationalized. This study proposes a transparent rule-based framework for estimating persistence, delegation-related lexical patterns, and refinement-related lexical markers without inferring psychological dependence. LMSYS-Chat-1M and WildChat were analyzed as primary conversational corpora, while Chatbot Arena was included as a structurally conditional A/B comparator. Across 1,879,085 valid analytical records, we calculated IPF, CDR, SRS, and the Conversational Behavioral Intensity Index (CBII), together with an equal-weight control. WildChat showed the highest mean CBII (0.2463), followed by LMSYS-Chat-1M (0.2123) and Chatbot Arena (0.1762). This ordering was stable under language controls, IPF thresholds of 5, 10, and 20 user turns, and question-level analysis of Chatbot Arena. Pairwise effect sizes were small to moderate, with the largest contrast between WildChat and Chatbot Arena (Cohen's
d
= 0.363).
Aracely Mera-Navarrete, Solange Revelo, Jefferson Beltrán-Morales et al.· Frontiers in Artificial Inte...· 0 citations
E EduFairBench provides a reproducible methodology for jointly analyzing predictive performance, robustness, uncertainty, and feedback quality, providing a comprehensive methodological framework for the rigorous evaluation of LLM-based educational assessment systems.
W. Villegas-Ch., Aracely Mera-Navarrete, Fernando Zúñiga-Tello et al.· Frontiers in Artificial Inte...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.