This work introduces Pluralis v0.1, a novel multimodal, multi-regional, and multilingual dataset built from a culture-first perspective and calls upon the research community to utilize this foundation to advance the science of multilingual, multicultural evaluation to better support AI cultural alignment globally.
Alicia Parrish, Rajat C. Shinde, Sanket Badhe et al.· arXiv.org· 0 citations
A unified systems foundation and reference architecture for the agentic skills ecosystem is established, formalize skills as externalized procedural knowledge bridging high-level cognitive planning with deterministic execution environments, and systematically delineate the architecture across a nine-stage lifecycle.
Sanket Badhe, D. Shah, Priyanka Tiwari et al.· 0 citations
SkillSec-Eval is presented, a lifecycle-aware framework for systematically evaluating the security of reusable agent skills and demonstrates that vulnerabilities arise at multiple lifecycle stages beyond execution, highlighting the need for lifecycle-aware security analysis of reusable agent skills.
The results demonstrate that explicit execution state is an effective and architecture-agnostic abstraction for scalable long-horizon agent skills, and improves task accuracy while substantially reducing cumulative token consumption.