Data availability is critical for understanding complex disease pathways and developing robust predictive models. Although high-throughput omics technologies have improved insight into disease mechanisms, data acquisition from inaccessible tissues such as the central nervous system remains a major limitation, causing small sample sizes and complicating early prediction of neurodegenerative disorders such as Alzheimer’s and Parkinson’s diseases. Generative modeling has emerged as a powerful approach for synthesizing data to support downstream clustering and prediction with small sample size, but existing methods rarely handle high-dimensional tabular omics data effectively. Adversarial random forests (ARFs) provide a well-performing framework for tabular data generation but are not designed for high-dimensional settings. To address this limitation, we introduce high-dimensional ARF (h-ARF), an extension of ARF optimized for integrated clinical and high-dimensional omics data. Using benchmarks across nine datasets and eight performance metrics, we show that h-ARF better preserves both feature distributions, and downstream clustering and prediction utilities compared with ARFs. The method is implemented in the opensource R package harf, available on CRAN.
C. Fouodo, J. Kapar, Anke Huels et al.· bioRxiv· 0 citations
Objective. The first objective of this study was to refine previously designed machine learning models that predict energy expenditure (EE) of preschool children by modifying the method used to calculate metabolic equivalents (METs). The secondary objective was to compare estimates of time spent in different physical activity intensities across the newly developed models, previously published METs models, existing METs-based models from the literature, and models calibrated using direct observation. Approach. The model training dataset included 35 Canadian children (aged 3.0–5.99 years) equipped with GT9X accelerometers on their right hip. A portable metabolic unit was used to measure EE during a semi-structured protocol consisting of activities ranging from low- to high-intensity. The resulting models were applied to a sample of Canadian preschool children (n = 118; aged 3.0–5.99 years) to estimate time spent in sedentary (SED), light (LPA), moderate-to-vigorous (MVPA), and total physical activity (TPA). A repeated measures ANOVA was used to compare time estimates across models and according to three different configurations of METs activity thresholds. Main results. Results indicated that the newly developed models from Objective 1 produced significantly different estimates of time spent in SED, LPA, MVPA, and TPA compared to both previously published models and other existing METs-based models, highlighting the impact of different approaches to calculating METs. Significance. Model selection and METs calculation methods markedly influenced activity intensity estimates, underscoring the need for consistent methodology. Classification models yielded the most plausible free-living estimates.
Hannah J. Coyle-Asbil, Katarina Osojnicki, Christoph Buck et al.· Physiological Measurement· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.