Skip to content
Preprint

L\'evy Attention: Single-Pass Predictive Uncertainty for Continuous-Time Attention

Aug 2026 · 0 citations · 43 references
Computer Science

TL;DR

This work shows the attention layer itself can close that gap between deep models for irregularly-sampled time series answer queries at arbitrary continuous timestamps: with the right stochastic formulation, the pass that makes each prediction also reports how far it should be trusted.

Abstract

Deep models for irregularly-sampled time series answer queries at arbitrary continuous timestamps, yet report nothing about how far each answer should be trusted. We show the attention layer itself can close that gap: with the right stochastic formulation, the pass that makes each prediction also reports, in closed form and at no extra cost, how far it should be trusted. We introduce L\'evy Attention, a cross-attention operator whose output is a stochastic integral against an inhomogeneous Poisson random measure: query-key compatibilities assemble an intensity over a continuous (time x channel) index space, the measure scatters atoms under it, and the output averages an interpolated value field at those atoms. In expectation it reduces to a mollified cosine-kernel attention, so it replaces a softmax layer and trains with exact gradients. What softmax discards, the Poisson construction preserves in closed form: the evidence $\Lambda_q$ (total compatibility mass) and the disagreement $\mathrm{tr}\,\Sigma_V(q)$ (value spread). An exact variance identity makes their combination $\hat\sigma(q)=\sqrt{\mathrm{tr}\,\Sigma_V(q)\,\varphi(\Lambda_q)}$ the root-mean-square deviation of the sampled operator, emitted by the deterministic pass with no trained head. Empirically, disagreement carries the signal, while the evidence factor swings from uninformative on dense data to strongly informative on sparse. On t-PatchGNN the operator swap costs at most 5.6% accuracy against a matched control and nothing on the sparsest dataset. The free disagreement signal improves on 20-pass MC dropout across matched five-seed suites, and $\hat\sigma$ scales a calibrated Gaussian whose zero-sample CRPS beats a fifty-draw sampler; a split-conformal wrapper reaches nominal coverage at every level, and one pass ranks 3,383 unseen patients by trust in 1.4 seconds.

View source

Similar papers

Preprint Aug 2026

Attention-Path Fragility as an Uncertainty Signal in Large Language Models

It is proposed that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways, a training-free estimator that masks attention heads and measures the BALD mutual information...

Minsoo Kim, Sungyoung Ji, Kisung Moon et al. · 0 citations
Preprint Aug 2026

Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No

FV-Action, the training-free method built on this analysis, is the strongest training-free result on this benchmark; it surpasses every TVG-trained model evaluated zero-shot on TACoS, and improves over direct prediction on ActivityNet Captions and QVHighlights, with no temporal supervision at any stage.

Ji Huang, Barry Devereux, Hui Wang · 0 citations
#machine learning Preprint Sep 2026

MemoryWalker: Stop Training Agents on Contexts They Never Saw

This work introduces two exact, gradient-equivalent corrections: LogitTree, a segmented K-forward traversal, and a packed 4D attention mask, and proposes SDCC (Self-Distillation for Conditioning Consistency), a single-backward-pass variational relaxation.

J. Zinco, Xun-Jie Zhu, Shen Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

RBS-Attention is proposed, a training-free sparse-prefill method with two complementary selection branches that controls the contribution of rescue blocks while preserving regular block-sparse FlashAttention execution and supports radius-adaptive dual-branch selection as an effective approach to long-context prefill.

Chu-Xu Song, Jiu-Qi Wei, Zhen-Can Peng · 0 citations
Preprint Aug 2026

TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity

We introduce TinyCast, an attention-free zero-shot forecaster that emits a predictive distribution from 146,505 parameters, on the premise that at this size the periodic structure of a context is worth computing rather than learning. A zero-parameter spectral detector supplies the dominant periods, the context is folde...

A. Steinhauser · 1 citation · ⚡1
Preprint Aug 2026

From Attention Masks to Inert Zero-Vector Tokens: OAttention and O-Closure for Token Dynamics

OAttention alone does not preserve a NULL state through ordinary host components, whereas the OTransformer path does, and a two-block ablation shows that OAttention alone does not preserve a NULL state through ordinary host components, whereas the OTransformer path does.

Heyang Gong · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.