AdaPCLA framework is proposed, which enables generative models to adaptively fit and generate EHR data through a data distribution-aware training strategy; this is achieved by internalizing data knowledge parameters by simulated annealing training.
Abstract
Generative modeling of longitudinal Electronic Health Records is increasingly important for privacy-preserving research, yet standard autoregressive models tend to underrepresent the co-occurrence structure of tail events (i.e., diseases, symptoms), reducing the fidelity and faithfulness of generated data for rare subpopulations. To this end, we propose AdaPCLA framework, which enables generative models to adaptively fit and generate EHR data through a data distribution-aware training strategy; this is achieved by internalizing data knowledge parameters by simulated annealing training. It also supports training-free adaptation to a diverse clinical population for generation through zero-shot distribution control. Moreover, our theoretical analysis characterizes rare-code logit updates through the label-wise empirical NTK and derives a prior-internalization bound for how annealing speed and NTK conditioning affect retained prior signals. Experiments on real-world data show that AdaPCLA achieves consistent gains in tail plausibility, downstream utility, and zero-shot control; in particular, it improves TailPairSeen over HALO by 114.2% on MIMIC-III and 65.1% on MIMIC-IV, outperforms GPT-style generation by 3.5% F1 for zero-shot cross-population adaptation.
Theoretically, it is proved that when the auxiliary outcomes satisfy a set of surrogacy conditions and the representation retains relevant covariate information, the original CATE is identified when the high-dimensional covariates are replaced by the learned representation.
Maitreyi Swaroop, Shikha Bhat, Samantha Rodriguez et al.· 0 citations
Synthetic data generation is increasingly used to enable data sharing and secondary analysis while protecting participant privacy, particularly for longitudinal tabular health data, where repeated measures per subject create within-subject dependence that most synthetic data methods are not designed to preserve. Existi...
Hanchang Cai, Wen-Shan Yu, Rui-Jin Lu et al.· bioRxiv· 0 citations
Self-supervised Causal Effects Estimation is proposed, a novel framework that integrates causal priors with self-supervised learning to construct balanced and predictive representations for causal effects estimation that consistently outperforms state-of-the-art methods.
Xin-Shu Li, Shiyi Yang, Venus Haghighi et al.· ACM Transactions on Intellig...· 0 citations
This work proposes generation-powered inference (GPI), a general framework for improving inference on distribution-valued parameters using auxiliary generative models, focusing on Wasserstein barycenters and related distributional functionals, and introduces a function-valued bridge representation.
This is the first work to formalise the balancing-prediction MI conflict and propose a structured resolution through complementary predictive and information-theoretic training objectives.
The combined objective is derived, which proof steps transfer from the classification setting without modification and which require adaptation, and which require adaptation on the PhysioNet 2019 Sepsis Challenge, treating the two hospital systems as sequential training fragments and the unseen system as an out-of-dist...
Behraj Khan, Shabir Ahmad, S. Bukhari et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.