A framework for AI-resilient assessment that shifts evaluation from product quality to demonstrable reasoning, decision-making, and ownership of learning is proposed, Illustrated primarily through health sciences education, with wider relevance to professional and practice-oriented disciplines.
Abstract
Generative AI tools can produce polished academic text on demand, undermining the validity of assessments that treat written submissions as evidence of individual learning. Detection-based countermeasures have demonstrated variable accuracy and equity concerns. This paper does not report empirical outcomes or validation data. It proposes a framework for AI-resilient assessment that shifts evaluation from product quality to demonstrable reasoning, decision-making, and ownership of learning. The framework comprises four pillars: (1) process-based documentation; (2) oral defense integration; (3) authentic task design; and (4) transparent AI-use policies aligned with intended learning outcomes. These are operationalized through a rubric model, an oral-defense protocol, and an assessment vulnerability audit tool. Illustrated primarily through health sciences education, with wider relevance to professional and practice-oriented disciplines, the framework treats generative AI as a catalyst for re-examining how assessment evidence is generated rather than solely as an integrity threat. All proposed tools are design instruments for future empirical evaluation, not psychometrically validated instruments.
This paper conceptualises the SAGE Defend step, the sixth stage of the Structured AI-Guided Education framework, as a format-agnostic assurance checkpoint for AI-integrated higher education assessment. The study responds to a verification gap identified in earlier SAGE research, in which process documentation and AI interaction logs were found to support transparency but not, by themselves, to verify individual ownership of reasoning in group-based AI-integrated submissions. Adopting a design-informed conceptual approach grounded in design-based research principles, the paper integrates a multi-year programme of empirical SAGE studies, a structured synthesis of the assurance-task literature, and diagnostic observations from three Defend-proximate assessment implementations across undergraduate and postgraduate units at Central Queensland University. It distinguishes between assurance tasks that directly require students to demonstrate reasoning or performance, controlled assurance conditions that restrict the assessment environment, and corroborative assurance signals that provide corroborating but non-stand-alone evidence. On this basis the paper proposes a three-class assurance-task typology, an epistemic matching framework, and six design principles for embedding SAGE Defend within assessment sequences. It further argues that assurance should be distributed across the assessment sequence of a unit, so that each learning outcome is verified at a point and intensity proportionate to its stakes rather than concentrated in a single terminal examination. The paper frames this response as assurance by design, an approach that, echoing the established engineering principles of security by design and privacy by design, builds verification into the assessment sequence rather than appending it after the fact, and it names the compounding cost of the retrofitted alternative as assurance debt. Rather than presenting SAGE Defend as an oral examination model or claiming empirical validation of a single format, the paper positions Defend as a design principle through which educators can align verification tasks with the cognitive, professional, or technical competency being assessed. The contribution is therefore conceptual and practice-informed, offering a structured basis for the future empirical validation of specific Defend formats across disciplines, cohorts, and delivery modes.
Mahmoud Elkhodr, E. Gide· Frontiers in Education· 0 citations
Assessment feedback on complex written reports remains one of the most persistent and resource-intensive challenges in higher education. However, no principled framework exists for deciding which feedback tasks might appropriately involve artificial intelligence and which must remain human responsibilities. This paper addresses that gap by proposing a tripartite feedback framework that distinguishes three analytically distinct levels: low-level structural and presentational feedback, intermediate-level factual content validation, and high-level critical evaluation and synthesis. Grounded in established feedback theory, including Hattie and Timperley’s feedback model and Boud and Molloy’s sustainable feedback design principles, the framework provides pedagogically justified criteria for allocating tasks between AI systems and human assessors, rather than automating whatever technology can technically perform. Five non-negotiable boundary principles govern any AI involvement at the intermediate level, preserving human oversight, academic accountability, and assessment integrity. This paper examines current technological capabilities and limitations at each level, proposes a phased implementation pathway with explicit human-in-the-loop requirements, and addresses implications for feedback literacy, student agency, equity, and security. A comprehensive mixed-methods evaluation design specifying the evidence required for empirical validation is also presented. The framework’s contribution lies not in prescriptive solutions but in providing structured categories, explicit boundary conditions, and validation criteria to guide context-sensitive institutional decision-making about AI integration in assessment.
This conceptual paper proposes a framework for redesigning asynchronous assessments in online graduate nursing education in response to generative artificial intelligence, and proposes four strategies grounded in the community of inquiry framework, authentic assessment theory, experiential learning theory, and transformative learning theory: experience-anchored analysis, multimodal documentation portfolios, asynchronous verbal exchange, and asynchronous collaborative process.
R. Anders, Kei-Shing Ng, Aderonke Odetayo et al.· Discover Education· 0 citations
The rapid diffusion of generative artificial intelligence (GenAI) tools such as ChatGPT has unsettled established academic practices related to assessment, authorship, integrity, and disciplinary knowledge production. This qualitative repeated cross-sectional study examines how faculty and staff at a large public research university understood these changes in 2023 and 2024. The analysis draws on responses to the same open-ended survey question collected from independent respondent groups in 2023 (n = 104) and 2024 (n = 313). Responses were analyzed inductively through thematic analysis and subsequently interpreted using Disruptive Innovation Theory and Complex Adaptive Systems Theory. The analysis identified both continuity and change across the two datasets. Responses in 2023 emphasized uncertainty, threats to academic integrity, and defensive assessment redesign. Responses in 2024 more frequently described pedagogical experimentation, process-oriented assessment, and AI literacy as emerging academic and professional competencies. Concerns about authorship, equity, reliability, and inconsistent institutional guidance persisted across both years. The findings suggest that faculty and staff discourse shifted from primarily containing GenAI-related risks toward selectively integrating the technology into teaching and professional practice. However, because the study used independent cross-sectional samples, it does not establish individual change over time. The study contributes a theoretically informed account of institutional sensemaking during the first two years following ChatGPT’s public release and identifies strategies for balancing innovation, integrity, equity, and the human purposes of higher education.
T. Balart, Gibin Raju, Kristi J. Shryock· Algorithms· 0 citations
This study investigated the impact of Department of Information and Communications Technology (DICT) CPD-accredited capacity-building programs on Artificial Intelligence (AI) in Region IX and BASULTA to develop evidence-based courses and digital competency advancement. As we navigate the complexities of the "age of intelligence" following the 2022 generative AI surge, this research explored participant proficiency across five specialized modules, while simultaneously documenting baseline implementation challenges and infrastructural friction. Utilizing a descriptive-quantitative research design, the study employed purposive sampling to evaluate training participants through the DICT Standardized Training Evaluation Tool. Key findings from empirical data gathered between March and May 2026 indicate exceptionally high proficiency levels among participants. Mean evaluation ratings consistently ranged between 4.0 and 5.0, specifically within the critical dimensions of Mastery of Subject Matter and Instructional Methodology. Furthermore, analysis indicated that the structured training successfully bridged demographic gaps. While demographic groupings showed no statistically significant variance in ultimate proficiency, systemic external challenges, most notably intermittent internet connectivity and course pacing, were identified as primary friction points affecting the learning experience. The study concludes that the CPD-accredited AI program effectively massified foundational AI literacy in historically underserved regions. The research recommends the institutionalization of a "Competency-Based Technology Skills Enhancement Workshop." This proposed framework focuses heavily on subject-specific AI applications and the integration of offline-compatible digital tools, thereby ensuring a resilient and sustainable pedagogical transformation in the current post-digital era.
Edilwaleed D. Hairon· International Journal of Res...· 0 citations
It is argued that detection-centred enforcement is a structurally weak control and proposed instead a layered institutional framework in which policy and governance, pedagogy and assessment redesign, and technology-based assurance operate as mutually reinforcing controls, sustained by a continuous audit and improvement cycle.
Dr. G. Purushothaman, Dr. S. Ganapathy, Mr. Saurabh Jaiswal, Mr. Thanga Kumaran M· International Journal of Adv...· 0 citations