TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation
Results show that execution evidence can expose behavioral failures missed by artifact inspection and can guide Skill generation toward jointly verified functional and safety outcomes.