Skip to content
Preprint

Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System

Jul 2026 · 0 citations
Computer Science

TL;DR

This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a graphical user interface (GUI) for visualizing outcomes.

Abstract

The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their response quality. Prevailing methods, often confined to singular dimensions, fall short of capturing the full spectrum of model capabilities. This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a graphical user interface (GUI) for visualizing outcomes. Evaluations on the TruthfulQA dataset unveil mainstream LLMs'strengths in reasoning tasks (peaking at a composite score of 0.6104) alongside pervasive limitations in navigating complex facts and ambiguities. Transcending the narrow lens of traditional metrics, this framework offers a transparent, adaptable avenue to illuminate model potential and deficiencies. Though presently focused on English tasks, its horizons beckon toward multilingual domains. This work carves a novel path for knowledge engineering and model refinement.

View source

Similar papers

Open access 2026

Challenging the Abilities of Large Language Models in Italian: a Community Initiative

The rapid progress of Large Language Models (LLMs) has transformed natural language processing and broadened its impact across research and society. Yet, systematic evaluation of these models, especially for languages beyond English, remains limited. "Challenging the Abilities of LAnguage Models in ITAlian" (CALAMITA) is a large-scale collaborative benchmarking initiative for Italian, coordinated under the Italian Association for Computational Linguistics. Unlike existing efforts that focus on leaderboards, CALAMITA foregrounds methodology: it federates more than 80 contributors from academia, industry, and the public sector to design, document, and evaluate a diverse collection of tasks, covering linguistic competence, commonsense reasoning, factual consistency, fairness, summarization, translation, and code generation. Through this process, we not only assembled a benchmark of over 20 tasks and almost 100 subtasks, but also established a centralized evaluation pipeline that supports heterogeneous datasets and metrics. We report results for four open-weight LLMs, highlighting systematic strengths and weaknesses across abilities, as well as challenges in task-specific evaluation. Beyond quantitative results, CALAMITA exposes methodological lessons: the necessity of fine-grained, task-representative metrics, the importance of harmonized pipelines, and the benefits and limitations of broad community engagement. CALAMITA is conceived as a rolling benchmark, enabling continuous integration of new tasks and models. This makes it both a resource – the most comprehensive and diverse benchmark for Italian to date – and a framework for sustainable, community-driven evaluation. We argue that this combination offers a blueprint for other languages and communities seeking inclusive and rigorous LLM evaluation practices.

Malvina Nissim, Danilo Croce, V. Patti et al. · 0 citations
Open access Jul 2026

Multi-Criteria Evaluation of Hierarchical Reasoning, Self-Correction, and Factual Consistency in Large Language Models across Complex Language Tasks

The rapid proliferation of large language models has necessitated the development of robust evaluation frameworks that extend beyond simple accuracy metrics. This paper introduces a comprehensive multi-criteria evaluation methodology designed to assess the capabilities of these advanced computational architectures in handling complex language tasks. Specifically, the study focuses on three foundational dimensions: hierarchical reasoning, self-correction mechanisms, and factual consistency. By systematically isolating these dimensions, the research provides a nuanced understanding of how models parse intricate problem structures, dynamically revise their internal states upon detecting errors, and maintain fidelity to established external knowledge bases. The proposed framework employs novel mathematical formulations to quantify these qualitative traits, enabling a rigorous, quantitative benchmarking process. Through extensive empirical analysis across diverse datasets, the findings reveal critical trade-offs between a model's ability to engage in deep hierarchical reasoning and its capacity to remain factually grounded. Furthermore, the evaluation of self-correction capabilities highlights persistent vulnerabilities in unsupervised revision protocols. This study contributes to the broader discourse on artificial intelligence reliability and safety by offering a structured approach to diagnosing model deficiencies, ultimately guiding the design of more resilient and dependable language processing systems

Stephanie Yam · 0 citations
Open access Jul 2026

MULTI-DIMENSIONAL TASK-ALIGNMENT FRAMEWORK FOR LARGE LANGUAGE MODELS: COMPARATIVE ANALYSIS OF ChatGPT, GEMINI, GROK AND CLAUDE

The proliferation of commercially available large language models (LLMs) has produced a competitive ecosystem in which model selection for specific professional tasks remains insufficiently theorized. These theses introduce the Multi- Dimensional Task-Alignment Framework (MTAF), a novel seven-criterion evaluation instrument designed to characterize the functional specialization of competing LLMs and translate benchmark performance into domain-specific selection guidance. Applying MTAF to four dominant systems – ChatGPT (OpenAI/GPT-4o), Gemini (Google DeepMind), Grok (xAI), and Claude (Anthropic) – we identify distinct competitive profiles: ChatGPT demonstrates leading performance in code generation, Gemini excels in multimodal and real-time grounded tasks, Grok provides unique access to temporally current social-media-derived data and Claude exhibits the highest reliability in long-document processing and complex instruction following. A derived task-model alignment matrix operationalizes these findings for practical decisionmaking across scientific research, software engineering, academic writing, and organizational management contexts.

Iryna Bobreshova, Olena Lebedieva · 0 citations
Preprint Jul 2026

From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models

A multi-layer taxonomy of 14 capability domains and 91 subskills across Primitive, Constructed, and Integrative layers is introduced and supports research organization, coverage audits, evaluation interpretation, and testable hypotheses for diagnosis, training, and transfer.

Shixin Fang, Jiachen Wo, Wenjuan Qin et al. · 0 citations
Preprint Aug 2026

MGAL: A Multilingual Granularity-Aware Long-Context Benchmark

Evaluation of long-context Large Language Models (LLMs) has advanced rapidly. However, most existing benchmarks are limited to the document level and focus mainly on high-resource languages, leaving many fine-grained challenges insufficiently evaluated. To address this gap, we present MGAL, the first multilingual, granularity- and position-aware long-context benchmark. MGAL is constructed from United Nations (UN) reports spanning 8K to 128K tokens across the six official UN languages. It covers four coherent levels of linguistic granularity (word, sentence, paragraph, and document) and further stratifies entries by their position within the document (begin, middle, and end), indexed at both the document and paragraph levels. This design enables systematic diagnosis of multilingual long-context comprehension across different granularities. Through extensive experiments and analyses, we find that: (1) LLMs perform well at word-level tasks but struggle with coarser-grained ones; and (2) Closed-source models retain a clear performance advantage in lower-resource languages. We further identify two new challenges: (1) Under local semantic crowding, where neighboring sentences share topics and entities, models tend to follow surface cues (e.g., connectives like ``however''or repeated entities) rather than the discourse role of the sentence in surrounding context (e.g., background, outcome); and (2) A gap between fluency and consistency in generated outputs, where models produce text that reads smoothly but drifts from the source facts. In addition, we observe several patterns in line with prior studies, including reliance on nearby evidence and reuse of options under uncertainty.

Chunhan Li, Chenglin Xu, Zongyang Zhang et al. · 0 citations

Have Large Language Models Improved Research Methodology?

Whether contemporary LLMs can reproduce the research outcomes of a fully documented human study: a 1991 article that identified dermatophytosis (ringworm) in historical fine art was evaluated.

Fredric Narcross, Robert Marks · 0 citations