We introduce a data-driven framework that maps noisy oscillatory time series directly onto the Hopf normal form, enabling inference of underlying dynamics without knowledge of governing equations. By embedding the normal form in a probabilistic state-space model, the method jointly infers latent states and system parameters, yielding robust estimates of the natural frequency, Floquet exponent, and asymptotic phase even far from the bifurcation point and under strong noise. Combined with complex Gaussian process regression, the approach further reconstructs phase and amplitude sensitivity functions from data. Benchmarks on the van der Pol oscillator demonstrate substantially improved accuracy and noise robustness compared with existing phase-based and regression methods. This work establishes a direct bridge between normal-form theory and statistical inference, providing a general and practical route to low-dimensional descriptions of oscillatory dynamics in complex systems.
We study the reconstruction of an unknown dynamical system from a single noisy scalar time series. The goal is to recover the underlying dynamics for forecasting. We introduce a method that uses differential embedding coordinates to identify a rational closure of the embedding dynamics directly from data. The closure is identified through a weak-form regression pipeline, which avoids unstable pointwise differentiation of noisy data. When applied to noise-free Lorenz and R\"ossler systems, the method recovers closures that support long forecasts across a broad ensemble of realizations ($18.1$ and $7.1$ Lyapunov times respectively). Under $15$--$30\%$ additive Gaussian noise, performance becomes system-dependent. For the Lorenz system, forecast horizons remain short even in the best cases, whereas the R\"ossler system generally performs better in absolute terms, though not once normalized by the Lyapunov time. Our proposed method recovers directly interpretable closure coefficients which we compared against the known analytic closures of the Lorenz and R\"ossler systems.
We present a machine learning framework for identifying sparse, interpretable models of dynamical systems directly from time-series data. Our approach parameterizes the underlying vector field using a neural architecture and trains it by minimizing a multi-step prediction loss over a finite horizon. To ensure numerical tractability, we optimize a mean absolute error objective averaged across prediction steps, and progressively increase the horizon during training. A key feature of this formulation is that it enforces consistency under repeated composition of the learned dynamics. As a result, the identified models exhibit significantly improved stability compared with approaches based on one-step regression of the vector field. When combined with sparsity-promoting regularization, this leads to parsimonious models that generalize beyond the training data. We demonstrate accurate recovery of systems exhibiting a wide range of behaviors, including stable and unstable fixed points, periodic orbits, and chaotic attractors. For chaotic systems, while long-term trajectory prediction is inherently limited by sensitivity to initial conditions, we show that multi-step training yields models with accurate short-term dynamics and strong agreement in long-time statistical properties, including mean, variance, and Lyapunov exponents. Moreover, we establish theoretical bounds linking trajectory error to statistical accuracy, providing a step toward a principled explanation for this behavior.
Identifying stochastic dynamical systems from observational data remains a major challenge in applied mathematics and engineering, particularly when complex systems are influenced by random perturbations and incomplete empirical information. This comprehensive review aims to examine state-of-the-art data-driven methods for discovering governing equations, estimating parameters, and predicting the behavior of stochastic dynamical systems. The review systematically analyzes key methodological approaches, including Sparse Identification of Nonlinear Dynamics (SINDy), Dynamic Mode Decomposition (DMD) and its extensions, Koopman operator theory, neural ordinary differential equations, and Bayesian inference. Each approach is evaluated in terms of its theoretical foundations, computational requirements, robustness to noise, and applicability to different classes of stochastic systems. Drawing on numerical experiments and real-world case studies, the findings show that no single method consistently outperforms others across all scenarios. Instead, hybrid approaches that integrate physics-informed constraints with machine learning demonstrate the strongest potential for advancing data-driven system identification. The review concludes that future research should address real-time identification, uncertainty quantification, and the integration of multi-fidelity data sources to improve the reliability and scalability of stochastic system modeling. This work contributes a comprehensive framework for guiding researchers and practitioners in selecting and implementing appropriate identification methods for stochastic dynamical systems.
Rishav Jha, Kameshwar Sahani, S. K. Sahani et al.· African Multidisciplinary Jo...· 0 citations
Functional dependence measures have become an important tool in the analysis of nonlinear time series and are typically formulated with respect to a given innovation representation of the process. This note points out that the probability space on which such representations yield the expected memory loss properties may not always coincide with the natural dynamical probability space of the model. We exhibit classes of uniformly ergodic autoregressive processes for which the behavior of the natural innovation coupling undergoes a qualitative transition as the model parameter varies. For this family of models, this transition coincides with a change in the sign of an associated Lyapunov exponent. In particular, a positive Lyapunov exponent may prevent the forgetting of initial perturbations along trajectories driven by the same innovations, despite uniform ergodicity of the associated Markov chain. These observations highlight the importance of carefully specifying the underlying probability space when interpreting or applying functional dependence measures.
It is demonstrated that the squared Pearson correlation coefficient provides a simple quantitative criterion for distinguishing chaos from noise directly from observed time-series data.
Across hydrodynamics, ecology, neuroscience, network dynamics, non-Hermitian physics, and socio-economic systems, asymptotically stable dynamics can exhibit large transient amplifications that are invisible to eigenvalue-based analyses. The mechanism is geometric rather than spectral: perturbations entering along one direction may be expressed transiently along another, allowing asymptotic decay to coexist with strong transient or noise-driven amplification. We introduce non-normal directional response inference, a data-driven method for detecting this geometry from multivariate time series when the governing operator is unknown. A local linear operator is estimated from sliding windows and projected onto the dominant two-dimensional input-response subspace. The reduced dynamics are summarized by the eigenvalue splitting $\Delta$, eigenvector non-orthogonality $K$, and the scale-free ratio $R=K/K_c(\Delta)$, where $K_c(\Delta)$ is the two-dimensional threshold for transient amplification. Controlled benchmarks show that the reduced geometry, particularly $R$, can be recovered from finite data even when the full high-dimensional operator is poorly estimated. Tests across sample size, dimension, training horizon, spectral structure, and non-stationarity confirm that the relevant response geometry requires far fewer observations than full-matrix recovery. Applied in moving windows to electrohysterogram, seizure EEG, freezing-of-gait, and unstable push-up inertial recordings, the method reveals systematic changes around known physiological or behavioral episodes through shifts in $R$, changes in $\Delta$, or stronger projection of fluctuations onto the inferred response direction. It thus exposes interpretable changes in local response geometry without framing the problem as supervised event detection.