Explainable and Risk Aware Autonomous Decision Systems A Scalable Secure by Design Architecture for Large Scale Distributed Computing Systems
This work introduces an explained, risk-conscious, and secure-by-design autonomic architecture of large scale distributed computing systems defining dynamic cloud-edge environments. The main issue it is expected to fix is the shortage of transparency, uncertainty quantification, and combined cybersecurity of the traditional black-box AI-driven distributed systems that restricts trust, scalability, and resiliency. In order to address these limitations, the suggested methodology combines some structured data preprocessing (Z-score filtering, entropy and variance extraction), probabilistic risk modeling, explainable AI based on SHAP, and Kubernetes KEIO-enabled adaptive orchestration into a single layered framework. Experimental analysis on cloud anomaly data sets shows good predictive accuracy of 91.1 percent, recall of 86 percent as well as precision of 79 percent as represented in the confusion matrix of 3333 true negative, 1000 true positive, 261 false positive, and 160 false negative. Risk scores are centred on 0.0, with most of the risk scores falling within the range of - 0.3 to 0.3, which is stable with respect to calibration. System availability indicates 94% of uptime, Mean Time to Recovery (MTTR) of 3.45 and 48% of scalability efficiency, which all indicate high uptime and efficient fault management. SHAP analysis reveals that power consumption (0.28) and memory usage (0.27) had the largest contributions to risks. In general, the architecture is able to implement transparent, resilient and scalable autonomous decision-making which is appropriate to mission critical distributed infrastructures.