Maria-X: A Multimodal Transformer Aware of Missing Data for Reliable Diagnosis and Prognosis of Alzheimer’s Disease
Abstract
Missing values are a persistent issue across real-world datasets, and they often undermine the reliability of models built on top of that data. Most existing solutions have been developed and tested under the MCAR assumption a condition that is uncommon in actual data collection settings and therefore leaves open questions about real-world performance. This work addresses that gap by extending an imputation approach originally validated under MCAR to instead operate under the more realistic Missing At Random (MAR) mechanism. The model's architecture is left unchanged from its original form, which allows any change in performance to be attributed specifically to the shift in missingness pattern rather than to structural modification. To evaluate this, synthetic MAR data was generated at several missing-data rates, and a series of experiments measured imputation accuracy and consistency at each rate using standard evaluation metrics. The resulting framework proves robust under MAR conditions, with particularly strong performance in cases where missingness is driven by observed variables a property that points toward practical real-world applicability. Taken together, this work extends the framework's usefulness beyond the MCAR setting and establishes a foundation for future work on more complex missingness mechanisms, including Missing Not At Random (MNAR).