Skip to content
Preprint

BTS-AgentBench: A Deterministic, Replayable Pipeline from Read-Only Telemetry Logs to Agent Benchmarks

Aug 2026 · 0 citations · 18 references
Computer Science

TL;DR

A telemetry-to-episode construction method instantiated as BTS-AgentBench is presented, which normalizes BTS metadata and raw histories into a read-only tool store, compiles static tasks with tool-derived gold answers and evidence, and lifts retained tasks into typed, bounded operator-facing episodes.

Abstract

Industrial sites contain large volumes of read-only telemetry, but few benchmarks specify how to compile these records into executable multi-turn agent tasks. We present a telemetry-to-episode construction method instantiated as BTS-AgentBench. The pipeline normalizes BTS metadata and raw histories into a read-only tool store, compiles static tasks with tool-derived gold answers and evidence, and lifts retained tasks into typed, bounded operator-facing episodes. The 532-row release adds clarification, goal revision, timestamp policy, quality-gated reporting, and evidence attribution while preserving the source computation and split. Coded contract preflight reports zero findings, and the construction-exclusion controller completes 0/532 rows. Two independent raw-to-episode builds match all 11 logical tool-store exports and reproduce the released 356/87/89 train/dev/test artifact exactly. Applying the shared construction path to XAI4HEAT produces 204 episodes; on its 41-row held-out test split, the controller completes 0 rows and the retained GPT-5.5 execution completes all 41. Code, artifacts, and replay reports are available at https://github.com/kjy7567/BTS-AgentBench.

View source

Similar papers

Preprint Aug 2026

Agent-Native Telemetry: Verifiable State-Delta Evidence for Autonomous Operations

Operational telemetry is predominantly engineered for human reading: systems repeatedly serialize verbose prose, static keys, and redundant context across billions of log lines. As autonomous AI agents become primary operational consumers, feeding them traditional logs wastes scarce context capacity parsing lexical syn...

Jun-Fei He, De-Ying Yu · 0 citations
#artificial intelligence Preprint Sep 2026

VST: Verifiable Structured Transport for Auditable Agent-to-Agent Alpha Discovery

Single-run agent-to-agent alpha discovery results are reported descriptively, gross of costs, and are explicit about their limits throughout; in particular they do not isolate the effect of the leap machinery from the inherited search substrate, which is left to future work.

Yu-Qi Li, Siyuan Liu, Bing-Jun Liu · 0 citations
Conference Jul 2026

Metamorphic Testing of Multi-Agent LLM Systems: A Trace-Based Behavioral Oracle Framework

Multi-agent systems built on large language models (LLMs) are increasingly deployed for complex tasks requiring autonomous planning, tool use, and inter-agent coordination. However, the non-deterministic nature of LLM outputs and the emergent behavior arising from agent interactions render traditional test oracles inef...

Gopalakrishnan Marimuthu · 0 citations
Preprint Aug 2026

Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay

A deterministic, zero-model pipeline is compiled into agent memory with a deterministic, zero-model pipeline that segments a local capture stream into typed activity frames, bounded episodes carrying application, site, timing, input volume, and evidence pointers back to the raw rows, with no model in the loop.

Nossa Iyamu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.