CoJEPA shows that combining objectives with complementary inductive biases can substitute for scale, encouraging future work to invest in smarter training objectives over ever-larger models.
Gabriel Meseguer-Brocal, Yuexuan Kong, Romain Hennequin· 0 citations
Deploying a new control policy for voltage control in active distribution grids requires evidence that physical limits will be satisfied before the policy is tested on the physical grid. This assessment is difficult for two reasons. First, simulations cannot capture every disturbance, modeling error, and device interaction present in the real grid. Second, historical measurements reflect operation under existing control policies, whereas a new policy may drive the grid into different operating conditions. To address these challenges, we propose Distributionally Robust Conformal Safety Screening (DR-CSS), a policy-agnostic framework for pre-deployment, scenario-by-scenario screening of a new control policy using historical data and a nominal simulator. For each new scenario, the simulator predicts a future voltage trajectory for the whole grid; DR-CSS then constructs a conformal safety interval around this prediction using historical simulation-to-reality errors. The interval is further enlarged to account for closed-loop changes induced by the deployment of the new policy and its interactions with the remaining controllers. To the best of our knowledge, DR-CSS is the first framework in power systems to combine historical data from an existing control policy with an imperfect simulator for pre-deployment safety screening of a new policy. Experiments on the IEEE 33-bus and IEEE 141-bus systems evaluate the deployment of learning-based voltage control policies and show that DR-CSS identifies all unsafe test scenarios. To reduce unnecessary warnings on safe scenarios, we adapt the safety intervals to different operating conditions and gradually introduce new policies with recalibration after each stage. These extensions increase the informational value of the safety screening and support safer deployment decisions in active distribution grids.
Sarra Bouchkati, P. Ellinas, Adriana Geisler et al.· 0 citations
This work introduces the first end-to-end neuromorphic spike-encoding and evaluation of the TIMIT dataset and quantifies the pipeline's efficiency with hardware-agnostic metrics based on the quantitative spiking activity.
Valentin Meunier, Amélie Gruel, Pierre Lewden et al.· 0 citations
SingProbe is introduced, a lightweight intrinsic runtime guard that directly reuses hidden states produced during LLM inference and operates alongside autoregressive decoding and extends this paradigm to medical generation through SingProbe-Med, which selectively activates risk-directed decoding interventions only when clinically relevant risks emerge.
Singg Team· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Results show that route-supervised frontier selection can improve budgeted search without altering biochemical generation, although performance remains dependent on frontier construction and reaction ranking.
Philippe Meyer, Guillaume Gricourt, T. Duigou et al.· 0 citations
MIOH provides a controlled framework for analyzing multi-image object hallucination and serves as a critical evaluation tool for developing more reliable multimodal AI systems.
Joonki Min, Chaeyun Kim, H. Choi et al.· 1 citation
In these experiments, BiG-SURE improves average abstention AUROC over prior black-box uncertainty estimators, while remaining simple, unsupervised, and applicable to black-box model settings.
It is found that training on the top 20% tokens ranked by GMTS consistently outperforms entropy-based token selection across three reasoning domains and various model sizes, suggesting that GMTS provides a more fine-grained estimate of token contribution for RLVR training.
This work investigates continued pre-training for adapting large language models to Swedish journalism, using a high-quality dataset that is curate from millions of news articles and demonstrates the importance of targeted evaluation in the adaptation process.
Lukas Borggren, Jenny Kunz, Marco Kuhlmann· 0 citations
The results indicate how trustworthy LLM-generated explanations are in model-free settings, where the same LLMs are used but no oracle exists to verify them.
The TSExplorer tool enables users to inspect high-dimensional datasets through multiple complementary 2D visualizations derived from high-dimensional feature representations to support a wide range of workflows.
Einari Vaaras, Manu Airaksinen, O. Räsänen· 0 citations
MineAmongUs is introduced, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action, and ARIA is proposed, a configurable VLM-agent harness that exposes five cognitive-component ablation axes and opens a new path for embodied VLM-agent alignment research.
Jaewoo Ahn, Junseo Kim, Hyunseo Kim et al.· 0 citations