This work presents an AI-assisted approach that generates candidate hazard scenarios from NASA's Aviation Safety Reporting System (ASRS), and proposes a hybrid variant, conditioning narrative generation on a structured hypothesis produced via evolutionary abduction, improving correctness and reducing variability.
Abstract
Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operations, and human factors - distinct from the functional hazard assessment applied at the aircraft-system level. We present an AI-assisted approach that generates candidate hazard scenarios from NASA's Aviation Safety Reporting System (ASRS). Given a target adverse outcome, it produces a structured hypothesis as categorical factors and a narrative scenario describing an operational event sequence consistent with the structure. Each scenario includes by a plausibility score from historical co-occurrence evidence and traceability to the most similar held-out ASRS reports. We then propose a hybrid variant, conditioning narrative generation on a structured hypothesis produced via evolutionary abduction, improving correctness and reducing variability. We evaluate multiple large language models, zero-shot versus few-shot prompting, and optional fine-tuning, measuring how prompting and model choice affect the validity and realism of the generated structures and narratives.
A safety-gated evaluation framework in which a trajectory succeeds only when all task goals are achieved without violating any hard safety constraint, while safe goal progress and trajectory safety are measured separately is established.
Yuchen Yuan, Zhenghuang Wu, Yuangan Li et al.· 0 citations
The reliability of aviation maintenance personnel directly impacts flight safety, yet systematic methodologies for the quantitative prediction of human error probability (HEP) in this domain remain lacking. To address this gap, a novel human factors reliability analysis method for aviation maintenance is proposed, extending the SPAR-H model through Evidential Reasoning (ER). This method is implemented as follows: Maintenance tasks are decomposed into subtasks. Subsequently, the eight types of Performance Shaping Factors (PSFs) for each subtask are evaluated by domain experts according to defined PSF levels. Expert judgments are then aggregated using Evidential Reasoning theory, enabling the calculation of aggregated PSF levels. These aggregated levels are interpolated to determine the corresponding impact multipliers. Finally, the HEP for aviation maintenance operations is calculated by integrating the SPAR-H basic error probability model with task series/parallel logic rules. The proposed methodology is validated using an inspection operation case study. This study establishes a methodological framework for human factors reliability analysis in aviation maintenance, providing a theoretical foundation for developing scientifically grounded prevention and control measures to enhance aviation safety levels.
Meng Meng, Ning Ma, Zhongqing Guan et al.· SAE technical paper series· 0 citations
Aviation safety improvement remains a continuous engineering challenge because modern aircraft are sensor-dense, software-intensive, and operationally interconnected, while health and operations data often remain fragmented across onboard, maintenance, dispatch, and ground systems. This fragmentation can delay weak-signal recognition and coordinated response. This paper proposes AI-SAFE FlightNet, a Safety-Aware, Intelligent, Federated, Explainable Flight Automation Network that integrates aircraft data intelligence, onboard edge AI, subsystem digital twins, phase-aware predictive analytics, evidence-grounded coordination, ground operations support, federated fleet learning, and governance by design. The framework supports pilots, engineers, dispatchers, and maintenance controllers without replacing certified human authority or certified aviation systems. A formal model combines sensor, flight-phase, environmental, and maintenance-history variables to estimate risk trajectories, prioritize alerts, and generate auditable response packages. The proposed evaluation uses public engine-degradation and trajectory datasets, synthetic subsystem telemetry, simulated maintenance records, and digital-twin experiments. Expected contributions are earlier warning, improved maintenance readiness, faster aircraft-to-ground coordination, privacy-preserving fleet learning, explainable recommendations, and measurable operational risk reduction. The objective is improved resilience and decision support, not zero risk or unrestricted autonomous control.
Prudvi Saisaran Ponduru, Pavani Priya Vyshnavi Nandanavanam, S. Ponduru· Engineering and Technology J...· 0 citations
Artificial Intelligence (AI) offers significant potential for future aviation systems; however, its integration into safety-critical applications requires compliance with the aviation sector's stringent safety standards. For AI and Machine Learning (ML)-based systems, the European Union Aviation Safety Agency (EASA) emphasizes the need to demonstrate the representativeness and completeness of the Operational Design Domain (ODD) and the associated data distributions used during development and verification. Despite this requirement, a structured engineering process for defining target distributions and evaluating representativeness within ODDs remains largely unexplored. This work presents a method for representativeness assessment of AI/ML constituent ODDs in the context of aviation safety assurance. Starting from the methodical identification of suitable target distributions, a process flow is proposed that guides developers from ODD definition and parameter distribution modeling to the quantitative assessment and interpretation of coverage results with respect to EASA's learning assurance objectives. As quantitative measures, the chi-squared goodness-of-fit test is examined and found unsuitable for the large data sets arising in this setting, leading to the adoption of the Kullback--Leibler divergence and Cram\'er's $V$ for the representativeness assessment. The method is demonstrated using the example of AI-based airborne collision avoidance, employing experimental data from previous Horizontal Collision Avoidance System (HCAS) and Vertical Collision Avoidance System (VCAS) simulations. The results illustrate how statistical distribution comparison methods can support the assessment of representativeness for safety-critical AI applications and contribute toward a systematic Safety-by-Design AI engineering process aligned with emerging EASA guidance.
Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann et al.· 0 citations
This study examined the prevalence and human factors correlates of startle and surprise events in commercial aviation using a large-scale analysis of NASA Aviation Safety Reporting System (ASRS) narratives. While startle and surprise effects have been identified as contributors to loss-of-control events and other critical incidents, prior empirical research has relied primarily on small-sample surveys, laboratory studies, or case analyses of individual accidents. The present study extends this work by analyzing the co-occurrence of pre-coded human factors with startle-related language across a decade of confidential incident reports. A keyword-matching algorithm was applied to 38,655 ASRS reports (2012–2022) to identify 2,642 reports (6.8%) containing startle/surprise-related language. Chi-square tests, odds ratio analyses with 95% confidence intervals, and logistic regression were used to compare human factors, flight phase distributions, anomaly types, and outcomes between startle-flagged and non-startle reports. All major human factors—including Fatigue (OR = 1.86, p < .001), Physiological factors (OR = 2.11, p < .001), Workload (OR = 1.60, p < .001), and Confusion (OR = 1.38, p < .001)—were significantly over-represented in startle reports. Logistic regression confirmed Physiological factors (β = 0.75) and Fatigue (β = 0.48) as the strongest independent predictors. Loss of aircraft control was 2.4 times more prevalent in startle reports. These findings provide large-scale empirical evidence that fatigue, physiological vulnerability, and high workload significantly amplify the risk and severity of startle reactions in operational aviation, supporting the development of targeted Crew Resource Management interventions and evidence-based training protocols.
The implementation of the ground deceleration function in civil aircraft represents a critically complex process that deeply relies on the seamless collaboration of multiple onboard systems, including but not limited to braking, thrust reversal, spoiler, and steering systems. The operational logic governing these systems is highly intricate, characterized by tightly coupled interactions, stringent safety requirements, and a vast array of diverse physical and logical interfaces. This inherent complexity makes it exceptionally difficult to gain a thorough, system-level understanding of the implementation mechanisms and collaborative principles solely through traditional means of examining extensive, yet often fragmented, design documentation. The limitations of document-based analysis frequently lead to unforeseen integration conflicts, which are typically discovered late in the development cycle, resulting in substantial rework costs and project delays. To address this pervasive industry challenge, this paper selects the aircraft ground deceleration function as a representative case study and proposes an innovative, simulation-based validation methodology. This approach systematically utilizes model state machines to create a dynamic digital representation of the system-of-systems, enabling rigorous validation of aircraft deceleration requirements under various operational scenarios. By adopting this model-based systems engineering (MBSE) paradigm for mechanism representation, our approach effectively captures the nuanced coordination, timing dependencies, and dynamic interactions within the multi-system operational logic. It thereby facilitates the intuitive identification, analysis, and resolution of potential design flaws, including logical conflicts, deadlocks, race conditions, and uncovered or ambiguous requirements. Consequently, the method not only provides a robust framework for validating the aircraft’s function-related design requirements with greater confidence but also offers crucial, data-driven support for the iterative optimization and evolution of the overall functional architecture. The fundamental value proposition of this research lies in its transformative capability to convert implicit design knowledge and assumptions—originally scattered across voluminous documents, specifications, and expert minds—into an integrated set of executable, observable, and analyzable formal models. This digital thread enables systems engineers and designers to identify deep-seated integration and coordination issues proactively during the early conceptual and detailed design stages, rather than relying on discovery during the late, costly integration and testing phases. By shifting validation left in the development V-cycle, this approach significantly reduces the risk of major design changes and associated cost overruns later in the project lifecycle. Ultimately, it effectively enhances the overall maturity, safety, certifiability, and operational reliability of complex aircraft function development, paving the way for more efficient and predictable engineering processes.
Mingqian Wang, Q. Yu, Miao Yu et al.· SAE technical paper series· 0 citations