Skip to content
Book Open access

CALM-TS: Risk-Controlled LLM Labeling for Time-Series via Calibrated Selective Gating

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 5674-5685 · 0 citations · 12 references

Abstract

Labeling time-series events such as anomalies and system failures is expensive and subjective, and distribution drift makes label definitions evolve over time. Large language models can generate weak labels with natural-language rationales, but their error rates are unbounded and their failure modes opaque. We present CALM-TS, a risk-controlled weak-supervision pipeline that bounds the labeling error rate while maximizing coverage. CALM-TS introduces behavioral probing over prompt surface form, sampling temperature, and temporal context window to elicit disagreement signals; a lightweight calibrator converts them into risk-bounded acceptance with finite-sample guarantees. CALM-TS attains 77--81% cost reduction on MIMIC-III and Yahoo~S5 at empirical risk ≤ a=0.05 in all 10 dataset-seed pairs. On the official PhysioNet Challenge 2015 binary alarm-verification task with gpt-4o-mini over five seeds, CALM-TS is the only method among nine LLM-as-weak-labeler baselines whose 5-seed mean risk strictly falls below the unfiltered LLM at non-trivial coverage, delivering a 15.0% relative reduction (0.400 → 0.340) at coverage 0.286. The framework yields auditable evidence chains and a coverage-based, label-free drift indicator.

Read PDF

Similar papers

Open access Jul 2026

ALSTMResNet: Active-Learning-Enhanced LSTMResNet as an Efficiency-Oriented Training Framework for Well Anomaly Monitoring Under Partial Labels

Oil-well anomaly monitoring supports safe and efficient oil-and-gas production, but delayed recognition of abnormal operating states can reduce lifting efficiency, trigger costly interventions, and increase operational risk. Existing data-driven detectors are also vulnerable to optimistic estimates when segmentation, n...

Feng Ge, Zhi Yang, Fan Yu et al. · 0 citations
Preprint Aug 2026

FETERS: Few-Shot Early Time-Series Classification via Effective Ratio Selection

Early time-series classification (ETSC) aims to make accurate predictions from partially observed time series as early as possible. Although various stopping mechanisms and feature learning strategies have been developed for ETSC, most existing methods assume access to sufficient labeled training data, which may be unr...

Chen-An Tai, Yujia Wu, Vincent S. Tseng · 0 citations
Preprint Aug 2026

Respect Your Zero-Shot Uncertainty: Conservative Calibration for Test-Time-Adapted Vision-Language Models

It is shown that TTA can increase confidence and reduce entropy even when the top-1 prediction and its correctness remain unchanged, a failure mode the authors term prediction-preserving sharpening, and proposed Zero-Shot-Anchored Entropy Calibration (ZAEC), a label-free post-hoc method that uses zero-shot entropy as a...

Jing-Yan Jiang, Yaru Sun, Xiao Chen et al. · 0 citations
Open access Aug 2026

Label space reduction for transductive zero-shot classification with large language models

This work proposes distilling the model into a probabilistic classifier, enabling lightweight deployment without repeated LLM calls, and demonstrates that LSR improves macro-F1 scores by an average of 7.0% compared to standard zero-shot classification baselines.

Nathan Vandemoortele, Bram Steenwinckel, F. Ongenae et al. · 0 citations
Preprint Sep 2026

When is Test-Time Adaptation Identifiable From Unlabeled Evidence?

Test-time adaptation (TTA) offers many ways to update a deployed model without labels, but choosing the wrong update can make a strong source model worse. Recent methods therefore try to predict which adaptation will work from unlabeled test data. We ask a prior question: does the evidence given to the selector contain...

Kartik Jhawar, Li-Po Wang · 0 citations
Jul 2026

Estimating Rare Events in Language Models with Proper Evaluation

This work introduces Gradient Activation Adaptive Multi-Level Splitting (GA-AMLS), which adapts rare-event Monte Carlo methods to the continuous activation space of language models and establishes activation space as a tractable domain for rare-event estimation in language models, circumventing the brittleness of discr...

Nikita Y. Parulekar, Anqi Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.