The Workflow Signal Protocol (WSP) is introduced, a deployment-layer measurement method for recording workflow observability as a structured workflow-observability record that supports the central claim that deployment-time workflow observability is measurable within the evaluated legal-AI settings.
Abstract
Legal AI benchmarks, citation checks, and retrieval-grounding tests primarily evaluate upstream capability: whether a model can answer, extract, or ground a legal task. Deployment asks a different question: whether a particular output remains observable enough to be deployed, reviewed, corrected, or escalated once it enters an organizational workflow. We introduce the Workflow Signal Protocol (WSP), a deployment-layer measurement method for recording workflow observability as a structured workflow-observability record. WSP encodes source status, proposition support, review state, recourse, provenance, and role-scoped disclosure. We validate WSP through controlled stress tests, public legal datasets, documented real-world failures, and a live-output pilot using three general-purpose model application programming interface (API) arms. In the main matched-vocabulary stress test, local formal/substantive routing reduced hidden-risk deployment from 93.5% under calibration-only abstention to 4.0% or below; all 192 pilot outputs were expressible as WSP records. These results support the central claim that deployment-time workflow observability is measurable within the evaluated legal-AI settings; validation in deployed institutional legal workflows remains future work.
APIFlow-Bench is introduced, a fully auditable benchmark for long-horizon, dependent REST-API workflows that decomposes performance into seven engineering capabilities and requires agents to produce answers supported by the actual call path.
Zelin Wan, Arash Nourian, Xiao-Xiao Li et al.· 0 citations
Large language model (LLM) agents can now carry out long-horizon technical workflows involving complex tool use, code execution, file edits, and generated artifacts. As agents do more work faster, the productivity bottleneck shifts from producing outputs to auditing whether those outputs are correct and trustworthy. Ag...
Daehong Kim, Hai-Chao Miao, Shu-Sen Liu· 5 citations
Enterprises deploy systems, not checkpoints. Usable capability depends jointly on weights, serving route, precision, output contract, and harness, yet all 18 audited benchmarks score advertised model identifiers. We treat this as measurement error and give a protocol that makes it reportable. It has three parts. A gold...
Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan· 0 citations
Science gateways often capture workflow execution but not the rationale behind human decisions such as dataset selection, algorithm choice, parameterization, and result interpretation, which limits reproducibility. We propose a framework that captures and shares scientific decisions as verifiable and reusable knowledge...
Pouriya Miri, V. Stankovski, D. Lavbič et al.· SN Computer Science· 0 citations
The results support verifier-selected supervision for bounded, machine-checkable capabilities, not arbitrary planning, enterprise validity, or unrestricted self-improvement.
WuYu-EnvLE-Bench is introduced, a benchmark built from real enforcement cases, regulatory standards, and expert review that highlights the need for evidence-grounded, rule-aware, and task-adaptive enforcement reasoning.
Zi-Liang Yang, Yi Zhang, K. Lin et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.