Data-driven methods have revolutionized ocean modeling, yet current approaches rely heavily on complete reanalysis datasets, imposing computational constraints and limiting model performance to that of the training data. Here, we present a generative state-space model and an optimization framework that enable learning directly from sparse and noisy observations. The model is essentially a hidden Markov model with a continuous state space, where oceanic physical quantities are treated as hidden states and measurements as observations, enabling a unified representation of ocean fields and observational data. Both the initial-state and state-transition modules are implemented as neural networks to capture the complexity and temporal evolution of ocean states, while the emission module is formulated as a masked Gaussian distribution. To train the model from sparse observations, we derive an optimization framework based on the expectation-maximization (EM) algorithm. The framework alternately reconstructs high-fidelity ocean fields via Langevin dynamics and optimizes deep neural networks to capture temporal evolution. Theoretical analysis shows that the framework maximizes the likelihood of observations under the generative model. For efficiency, we assume that ocean-state evolution follows a stationary, ergodic, and Markovian stochastic process and adopt only length-two state sequences during optimization. Experiments on CMIP6 simulation data and FY-3D satellite data demonstrate high-fidelity reconstruction and accurate prediction, showing that sparse observations can directly improve the model's representation of ocean-state dynamics. This work offers a scalable pathway for next-generation Earth system models to learn directly from sparse, incomplete real-world observations.
Yangyang Kong, Yutong Jiang, Yanhai Gan et al.· 0 citations
ABSTRACT Haploid induction coupled with genome editing (HI‐Edit) enables direct modification of commercial crop varieties, bypassing the need for trait introgression or direct transformation of elite lines with CRISPR machinery. However, its widespread application has been constrained by low haploid editing rates (HER), the proportion of haploids carrying edits within the short window between double fertilization and uniparental chromosome elimination. Here, we report substantial improvements in maize HI‐Edit efficiency through three complementary strategies: (1) driving an optimized LbCas12a variant (LbCas12aV) using promoters that are highly active in sperm cells and early zygotes; (2) applying a post‐pollination heat treatment; and (3) fusing LbCas12aV with the UBA2 domain (ubiquitin‐associated domain‐2 of Arabidopsis thaliana RAD23) to enhance protein stability during haploid induction. Post‐pollination heat treatment alone increased HER to 19.1% (up to 12‐fold improvement depending on the target site), providing a simple and effective method to boost the yield of edited doubled haploid (DH) plants. UBA2 fusion improved HER by 6‐fold at the Waxy1 (Wx1) locus and 4.5‐fold at the Glossy2 (Gl2) locus under normal conditions. Strikingly, combining UBA2 fusion with heat treatment raised the average HER to 25% across multiple events targeting Wx1, with the highest HER reaching 33%. Collectively, these findings demonstrate that increasing CRISPR‐Cas protein abundance and modulating environmental conditions can overcome key bottlenecks in HI‐Edit. We establish a robust, scalable framework that is readily transferable to other crops for elite‐line genome editing.
Dawei Liang, Huanhuan Guo, Juan Wei et al.· Plant Biotechnology Journal· 0 citations