2026· Informatization and communication· Vol 2· 0 citations
TL;DR
The aim is to develop the architecture of a two-level evaluation system featuring an Evaluator Agent with reasoning chain tracing, a control dataset of 30+ use cases, and a validation mechanism on samples of real user queries.
Abstract
The article addresses the problem of evaluating the quality of large language models (LLMs) and intelligent agents in multi-agent competitive intelligence automation platforms. The aim is to develop the architecture of a two-level evaluation system featuring an Evaluator Agent with reasoning chain tracing, a control dataset of 30+ use cases, and a validation mechanism on samples of real user queries. The proposed architecture includes Cell Agent, Cell Agent Models, Evaluator Agent, and a results database. A four-criteria verification model is introduced: (a) step verifiability and logical consistency, (b) source correctness and relevance, (c) result correctness, and (d) solution path optimality. The scientific novelty lies in substantiating an integrated evaluation approach unifying model benchmarking and agent tracing within a single architecture.
The Agent-to-Agent (A2A) Protocol was developed and released as an open-standard, vendor-agnostic, AIinteroperable agent solution through Google, then transitioned to the Linux Foundation in April 2025. It was the first open standard to be developed to support A2A interoperable agents. Its primary design tenet is the a...
Sunera Godakanda, K. Vidanage, Sabyasachi Bhattacharyya et al.· International Conference on...· 0 citations
AgentFactory is presented, a framework that jointly optimizes both foundation models and workflow structures in agentic systems while considering multiple objectives including performance, cost, and efficiency, and establishes AgentFactory as a promising approach for developing more capable and efficient agentic system...
En-Ci Zhang, Hao-Fen Wang, Yue-Sheng Zhu et al.· Pacific Rim International Co...· 0 citations
A unified, taxonomy-driven, and deployment-oriented survey of agentic AI systems, synthesizing recent advances through a modular reference architecture and a four-dimensional taxonomy that characterizes agents along the axes of autonomy, tool use, collaboration, and safety–governance is presented.
Sparsh Bajoria, Shreyanshu Ranjan, Adhitya M et al.· Cognitive Computation· 0 citations
This paper proposes the Evaluation Context Protocol (ECP), an early-stage, vendor-neutral framework intended to act as a portable evaluation contract layer for agentic systems and describes an open-source reference implementation that includes adapters for LangChain, LlamaIndex, CrewAI, and PydanticAI.
A comprehensive overview of the existing tools and frameworks for implementing MAS in software engineering and a set of lessons learned and challenges that can help researchers and practitioners to select a suitable MAS framework according to their needs are provided.
Maria Sâmyla Serafim de Oliveira, M. Ibiyo, Marco Gianrusso et al.· 1 citation
A comprehensive framework for the design, evaluation, and responsible deployment of Agentic AI is proposed, emphasizing safety, explainability, human-in-the-loop supervision, and ethical compliance and aims to maximize the benefits of Agentic AI while minimizing potential risks.
Nitin S. Shrirao, Dnyaneshwar S. Jadhav, Sarita B. Patil· Recent Trends in Mathematics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.