Large Language Model (LLM) agent frameworks such as LangChain, LlamaIndex, and CrewAI have become critical infrastructure powering production AI systems, yet they remain severely under-tested due to fundamental challenges in automated testing. Unlike traditional software, where crashes serve as reliable oracles, defect...
LeanGuard is presented, a neuro-symbolic framework that assigns each act to the side equipped for it, and it is argued that the remedy is not better prompting but a separation of roles: the component that interprets the code must not also be the one that decides a safety obligation is met.
Yanjie Zhao, Hongjie Chen, Li Lu et al.· arXiv.org· 0 citations
IAL-Scan is proposed, a static analysis tool for detecting IAL failures in real-world LLM agent projects, and builds an Agentic Loop Dependence Graph (ALDG) to recover explicit and framework induced feedback paths, and checks whether these paths can repeatedly reach costly or state growing operations without an effecti...
Xinyi Hou, Shenao Wang, Yanjie Zhao et al.· 2 citations
The evaluation shows that AgentFlow recovers richer agent entities and dependencies than existing AST-based agent static analysis tools, generates more dependency-aware Agent BOMs, and uncovers 238 taint-style prompt-to-tool risks in real-world agent programs.
Shenao Wang, Xinyi Hou, Yanjie Zhao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.