Skip to content
Preprint

WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management

Jul 2026 · 0 citations · 20 references
Computer Science

TL;DR

WuYuEval is introduced, a multi-level benchmark for evaluating LLMs in SWM across foundational knowledge, domain reasoning, and expert decision-making and suggests that visible deliberation helps only when it remains anchored to units, assumptions, and engineering constraints; otherwise, it may drift from decisive answer boundaries.

Abstract

Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to assess because existing benchmarks emphasize general knowledge rather than professional decisions under engineering, environmental, and policy constraints. We introduce WuYuEval, a multi-level benchmark for evaluating LLMs in SWM across foundational knowledge, domain reasoning, and expert decision-making. After quality auditing, WuYuEval contains a Foundation Module with 4,590 closed-ended multiple-choice questions across six task types and eight domain categories, together with an Expert Module with 247 scenario-based open-ended questions involving multi-objective optimization, constraint trade-offs, and system design. For expert tasks, we combine anchor-calibrated LLM-as-a-Judge scoring with Elo-based pairwise comparison. Across 33 LLMs, performance varied widely. The leading model reached 94.64\% accuracy on the Foundation Module, but average accuracy still fell from 84.14\% on easy questions to 42.50\% on hard questions, with lower performance concentrated in calculation, experimental design, urban planning, and open-ended expert tasks. Reasoning-oriented Thinking modes improve most matched model pairs after auditing, but the gains depend on baseline capability and are not uniformly positive. These results suggest that visible deliberation helps only when it remains anchored to units, assumptions, and engineering constraints; otherwise, it may drift from decisive answer boundaries. WuYuEval therefore provides both an evaluation resource and an empirical basis for developing SWM-oriented foundation models with professional reasoning chains and explicit constraint control.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial Scenarios

Large Multimodal Models (LMMs) large-scale deployment in industrial warehouse settings specifically necessitates that models exhibit human-expert-level hazard-oriented perception, understanding, and reasoning capabilities. However, the scarcity of real industrial data, tightly coupled to commercial terms, significantly...

Hanjing Zhou, Mingze Yin, Ying Lian et al. · 0 citations
Preprint Aug 2026

BC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERP

BC-Bench is introduced, a benchmark designed to evaluate agentic engineering on real-world tasks in AL, the DSL for Microsoft Dynamics 365 Business Central, and evaluates multiple frontier models across two agent harnesses, utilizing multi-run metrics to account for nondeterminism.

Haoran Sun, K. M. Hansen · 0 citations
Review Aug 2026

FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review

FinRiskAtlas is introduced, a Chinese-language benchmark that evaluates financial LLMs along two complementary dimensions: operation execution under fixed evidence states and evidence-state control under evolving review conditions, and shows that broad financial capability scores do not fully capture where models are r...

Su-Yang Zhong, Jingzhe Zhu, Qi Xu et al. · 0 citations
#machine learning Preprint Sep 2026

Evaluating Open-Weight LLMs for Turkish Domain Documents Under Retrieval and Hardware Constraints

Most Turkish-capable large language models (LLMs) are evaluated using general-purpose benchmarks rather than long, structurally complex domain documents. This paper evaluates five open-weight 7B-8B models for Turkish document question answering under a resource-constrained local deployment setting. The primary benchmar...

Imtiaz Ul Hassan, Öykü Akbulut, Onur Kaya et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories

Large language model based coding agents have made substantial progress on repository-level software engineering tasks. Existing repository benchmarks, however, usually start from a human-identified issue and evaluate whether a patch satisfies a functional signal. We present SWE-Prometheus, a benchmark for the broader...

Jia-Jun Wu, Lei-Xin Sun, Zi-Hang Tan et al. · 0 citations
Review Open access Sep 2026

Performance and Consistency of Large Language Models in Key Labor-Intensive Tasks of Systematic Reviews.

OBJECTIVE To evaluate the performance and consistency of Large Language Models (LLMs) in core systematic review (SR) tasks and to introduce open-source tools for automated batch processing that provide decision rationales. METHODS We assessed GPT-4o, Kimi-K2, DeepSeek-V3, and DeepSeek-R1 on five SR tasks: title/abstr...

Yi-Ran Liu, Xi-Ling Wang, Zi-Xuan Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.