Predicting whether an LLM agent will fail has emerged as a promising direction for supporting intervention during execution. Recent approaches report strong predictive performance, often with AUROC values between 0.85 and 0.94. However, predictors are typically trained by pooling runs from many tasks. We hypothesize th...
Mohsen EsfandyariDoulabi, L. Arkoh, B. Tadesse et al.· 0 citations
PyMOP is presented, a generic, extensible, and more efficient Python RV tool that supports five logics, implements five monitoring algorithms, ships with 81 specs, and supports three instrumentation strategies that help find hundreds of bugs by monitoring test executions against formal specifications.
Zhuohan Shen, Mohammed Yaseen, Kevin Guan et al.· SIGSOFT FSE Companion· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.