This work proposes GAD-MoRE, a novel framework for zero-shot Generalizable Graph Anomaly Detection with a Mixture of Riemannian Experts architecture, which significantly outperforms state-of-the-art generalist GAD baselines in the zero-shot setting.
Xinyu Zhao, Qingyun Sun, Jiayi Luo et al.· arXiv.org· 0 citations
Experiments on small-scale datasets with simulated ground truth across the full data distribution show consistent accuracy gains over baselines, demonstrating the method's effectiveness in data-scarce manufacturing environments.
Dennis Gross, Helge Spieker, Arnaud Gotlieb et al.· arXiv.org· 0 citations
STGAT (Spatio-Temporal Graph Attention Network), a clock-dynamics-aware anomaly detection solution that jointly models temporal distortion and inter-device consistency in energy IoT systems, is introduced.
Saeid Jamshidi, Omar Abdul Wahab, Rolando Herrero et al.· IEEE Internet of Things Jour...· 1 citation
Reach audiences
Advertise in front of researchers, engineers, and readers.
It is suggested that diffusion post-training selectively preserves or reorganizes inherited computation according to task structure, rather than uniformly replacing autoregressive mechanisms.
This paper formalizes KV management as a causal system of three primitives: KV Admission, Selection, and Eviction, and instantiate KV Admission via Write-Gated KV (WG-KV), a lightweight mechanism that learns to predict token utility before cache entry.
Yen-Chieh Huang, Rui Fang, Ming-Syan Chen et al.· 2 citations
Kascade is a training-free sparse attention method that leverages known observations such as 1) post-softmax attention is intrinsically sparse, and 2) the identity of high-weight keys is stable across nearby layers to achieve high accuracy on long-context LLM inference.
It is found that explicit world-modeling yields better representations in terms of higher probing accuracy and steerability of the model, and that better representations yield larger gains from GRPO, especially on harder cube states.
Prakhar Gupta, Henry Conklin, Sarah-Jane Leslie et al.· arXiv.org· 3 citations
ScalePRM, which scales verification compute as an alternative to ground-truth supervision for training process reward models, generates multiple independent verifications of each reasoning step and aggregate their judgments to produce synthetic step-level labels without ground truth.
Salman Rahman, Sruthi Gorantla, Arpit Gupta et al.· 0 citations
A curiosity-driven quantized Mixture-of-Experts framework that addresses both accuracy and stability through Bayesian epistemic uncertainty-based routing across heterogeneous experts, suitable for safety-sensitive edge deployments where both accuracy and predictability are critical.
S. C. Cajas Ordóñez, Luis Fernando Torres Torres, M. J. Meni et al.· arXiv.org· 1 citation
A cross-fidelity knowledge distillation and adaptive fusion network (CFKD-AFN), which leverages abundant but low-fidelity simulation data to enhance the prediction on scarce but high-fidelity trial data, and is extended to an interpretable variant for exploratory analysis of feature-attribution patterns associated with treatment outcomes.
Wen-Jing Chen, Lian-Sheng Zhuang, Zi-Ying Luo et al.· arXiv.org· 0 citations
Recent advances in self-supervised learning for EEG representation have largely relied on masked reconstruction, where models are trained to recover randomly masked signal segments. While effective at modeling local dependencies, the training objective of masked reconstruction does not compel the model to capture global generative constraints essential for characterizing neural activity. To address this limitation, we propose EEGDM, a novel self-supervised framework that leverages latent diffusion models to generate EEG signals as an objective. Unlike masked reconstruction, diffusion-based generation progressively denoises signals from noise to realism, compelling the model to capture holistic temporal patterns and cross-channel relationships. Specifically, EEGDM incorporates an EEG encoder that distills raw signals and their channel augmentations into a compact representation, which serves as conditional information to guide the diffusion denoising process, thereby enabling the encoder and diffusion model to be jointly optimized through the generative objective. This design endows EEGDM with a compact latent space, which not only offers ample control over the generative process but also can be leveraged for downstream tasks. Experimental results show that EEGDM (1) reconstructs high-quality EEG signals, (2) learns robust representations, and (3) achieves competitive performance across diverse downstream tasks, thus exploring a new direction for self-supervised EEG representation learning.
Shaocong Wang, Tong Liu, Ming Li et al.· arXiv.org· 2 citations