Skip to content

Improving observability and reliability in SKA control software through TANGO-controls OpenTelemetry and CSP.LMC test analytics

Aug 2026 · Astronomical Telescopes + Instrumentation · Vol 14155, pp. 141553T - 141553T-9 · 1 citation · 11 references
Engineering

TL;DR

This work investigates how the integration of Tango OpenTelemetry instrumentation with CSP.LMC’s test-result collection framework can provide deeper insight into software behaviour under anomalous conditions and explores how this combined approach can support more rigorous performance benchmarking and contribute to higher software quality and system reliability across the SKA ecosystem.

Abstract

In a large-scale scientific project such as the Square Kilometre Array (SKA), reliability is critical. The ability to detect rare, hidden, and difficult-to-reproduce anomalies is essential and must be done as early as possible during development, well before integration and testing, commissioning, or production. In this context, robust Continuous Integration Continuous Delivery (CI/CD) practices and comprehensive software testing are mandatory requirements for all components. SKA Software Components rely on TANGO-Controls and run as containerised applications orchestrated via Kubernetes. Due to the complexity of the applications sometimes unpredictable behaviours can be observed and non-deterministic failures that must be carefully monitored and understood. To address these challenges, the Central Signal Processor Local Monitoring and Control software subsystem (CSP.LMC) has started a data-collection campaign aimed at gathering statistical information on test failures while also enabling systematic benchmarking of subsystem functionality. With the introduction of OpenTelemetry support in TANGO-Controls starting from version 10, new opportunities have emerged to significantly enhance observability within Tango Devices. This work investigates how the integration of Tango OpenTelemetry instrumentation with CSP.LMC’s test-result collection framework can provide deeper insight into software behaviour under anomalous conditions. Additionally, it explores how this combined approach can support more rigorous performance benchmarking and contribute to higher software quality and system reliability across the SKA ecosystem.

View source

Similar papers

Conference Aug 2026

Integrating Observability, DevSecOps, and Policy-as-Code for Enterprise Cloud Reliability

Modern enterprise cloud infrastructures need to be able to deliver software quickly, remain operationally resilient and follow strict regulatory requirements. Through traditional operations, observability, security and governance are typically treated as separate processes, which leads to longer Mean Time To Repair (MT...

Vishwa Lakhnakiya · 0 citations
Aug 2026

Making sense of metrics: monitoring the performance of the Rubin Observatory control system

The Rubin Observatory Control System coordinates several distributed components. Observability platforms have been set up to provide raw metrics, dashboards, and logs that describe this environment, but turning that information into operational understanding requires constant monitoring and analysis. This work focuses...

Amanda Ibsen, Michael Reuter, A. Fausti et al. · 3 citations
#small language model Open access Aug 2026

Let’s read the log: root cause analysis of railway test execution logs with large language models

Results showed that long-context LLMs tended to achieve higher accuracy than smaller models, suggesting that LLMs are currently better suited to support human-in-the-loop root cause analysis than to fully automate it, and motivating further work to improve prediction accuracy for log-based RCA.

Rahmanu Hermawan, Alessio Bucaioni, Eduard Paul Enoiu et al. · 0 citations
Open access 2026

Sustainable Compliance in Regulated Manufacturing: Deploying a Low-Code, AI-Driven eDHR Framework for Scalable Traceability

The deployment of a validated, low-code electronic Device History Record system on a high-mix, low-volume manufacturing line at Smith & Nephew’s Memphis site provides a practical and replicable model for regulated manufacturing environments that need to strengthen compliance while gaining flexibility for future analyti...

Pareshkumar Hotchandani, Mr. Wakhare · 0 citations
Open access Sep 2026

Resilient Test Automation Framework for Microservices: Addressing Integration, Scalability, and Reliability

A Resilient Test Automation Framework (RTAF) is proposed to address three critical dimensions of microservices testing: integration, scalability, and reliability to reduce dependency-related test failures, improve the detection of service incompatibilities, and enable systematic verification of system behaviour under c...

Awonowo Olusegun Oriyomi · 0 citations
Preprint Aug 2026

JTA: Joint Testability Architecture for Scenario-Based Validation of Safety-Critical Software

JTA provides an architectural basis for modeling, designing, and assessing scenario-based validation in safety-critical software and shows that link-loss scenarios are comparatively mature, whereas state-estimation anomaly scenarios remain harder to validate because evidence alignment and attribution semantics are weak...

Wenyao Xue, Jiandi Wang, Yichen Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.