Skip to content

Making sense of metrics: monitoring the performance of the Rubin Observatory control system

Aug 2026 · Astronomical Telescopes + Instrumentation · Vol 14155, pp. 141551C - 141551C-13 · 3 citations · 9 references
Engineering

Abstract

The Rubin Observatory Control System coordinates several distributed components. Observability platforms have been set up to provide raw metrics, dashboards, and logs that describe this environment, but turning that information into operational understanding requires constant monitoring and analysis. This work focuses on how the observatory software team monitors the health and performance of these interconnected components during operations. We present the metrics that have proven most effective for diagnosing issues and show how cross-system dashboards and targeted alerts enable rapid fault localization and resolution. We discuss real life examples, including the detection of broker-level issues, the correlation of subsystem glitches with resource saturation, and how monitoring helps us differentiate between localized component failures and issues broader in scope. Rather than detailing the characteristics of the observability stack itself, this work concentrates on practical application: how metrics guide investigation, how interpretations evolve with experience, and how this accumulated experience ultimately improves the stability and reliability of one of the most complex real-time telescope control systems in operation today.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.