Can Open-Weight LLMs Predict Test Outcomes from Traces? A Cross-Subject Boundary Study on Tests4Py
Determining whether a test execution should pass or fail remains a core obstacle to end-to-end test automation. Recent work has focused mainly on generating assertions from source context, leav- ing open how much pass/fail signal can be recovered directly from runtime behavior, especially under subject transfer. We pre...