Skip to content
Review Open access

From Layperson to Medical Expert: A Corpus-Based Analysis of Linguistic Variation in Audience-Conditioned Biomedical Discourse

Sep 2026 · TTU Journal of Science · 0 citations · 14 references

Abstract

Biomedical text accessibility is often evaluated through readability formulas, although audience adaptation may involve changes across multiple linguistic dimensions. This study analyzed the Persona-Guided Controllable Biomedical Summarization (PERCS) corpus, an expert-curated, GPT-4-assisted resource in which persona-specific drafts were reviewed and revised by physicians. Following corpus-integrity screening, the analytic dataset comprised 1,984 English-language summaries derived from 496 biomedical source abstracts, with matched versions for laypersons, premedical students, non-medical researchers, and medical experts. Linear mixed-effects models accounted for the repeated-source structure and tested categorical audience differences, ordered trends, and nonlinear trajectories[cite: 11]. Twelve theory-driven linguistic features were examined, with Flesch Reading Ease and Flesch–Kincaid Grade Level used as conventional readability baselines[cite: 11]. All 12 linguistic features showed significant overall audience effects after false-discovery-rate correction[cite: 11]. Information-packaging measures showed the clearest expertise-related progression, whereas several grammatical and cohesive features followed nonlinear or non-monotonic trajectories[cite: 11]. Differences between the premedical and non-medical researcher conditions were consistently smaller than those involving the layperson or medical-expert endpoints[cite: 11]. In a source-separated held-out evaluation, combining readability with the 12 linguistic features increased macro-F1 from 0.610 to 0.677 and reduced multiclass log loss from 0.804 to 0.664 relative to readability alone[cite: 11]. The advantage remained under length-controlled sensitivity analyses[cite: 11]. These findings indicate that audience-conditioned biomedical rewriting in PERCS is distributed across multiple linguistic domains and is uneven rather than reducible to a single readability gradient[cite: 11]. Because PERCS is a controlled, persona-conditioned corpus, the observed linguistic profiles should not be interpreted as evidence of actual reader comprehension or as equivalent to naturally occurring human audience accommodation[cite: 11].

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.