WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management

WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management

WuYuEval:面向固体废物管理领域大语言模型的多层次基准测试

Abstract: Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to assess because existing benchmarks emphasize general knowledge rather than professional decisions under engineering, environmental, and policy constraints.

摘要: 大语言模型(LLMs)正日益被用作技术助手,但由于现有的基准测试侧重于通用知识,而非工程、环境和政策约束下的专业决策,因此评估其在固体废物管理(SWM)领域的胜任能力仍然十分困难。

We introduce WuYuEval, a multi-level benchmark for evaluating LLMs in SWM across foundational knowledge, domain reasoning, and expert decision-making. After quality auditing, WuYuEval contains a Foundation Module with 4,590 closed-ended multiple-choice questions across six task types and eight domain categories, together with an Expert Module with 247 scenario-based open-ended questions involving multi-objective optimization, constraint trade-offs, and system design.

我们推出了 WuYuEval,这是一个用于评估大语言模型在固体废物管理领域表现的多层次基准测试,涵盖了基础知识、领域推理和专家决策三个维度。经过质量审核,WuYuEval 包含一个基础模块,其中有 4,590 道涵盖六种任务类型和八个领域类别的封闭式选择题;此外还有一个专家模块,包含 247 道基于场景的开放式问题,涉及多目标优化、约束权衡和系统设计。

For expert tasks, we combine anchor-calibrated LLM-as-a-Judge scoring with Elo-based pairwise comparison. Across 33 LLMs, performance varied widely. The leading model reached 94.64% accuracy on the Foundation Module, but average accuracy still fell from 84.14% on easy questions to 42.50% on hard questions, with lower performance concentrated in calculation, experimental design, urban planning, and open-ended expert tasks.

针对专家任务,我们结合了基于锚点校准的“大模型作为裁判”(LLM-as-a-Judge)评分机制与基于 Elo 等级的两两比较法。在 33 个大语言模型的测试中,表现差异显著。领先模型在基础模块上达到了 94.64% 的准确率,但平均准确率仍从简单问题的 84.14% 下降至困难问题的 42.50%,且在计算、实验设计、城市规划和开放式专家任务中表现较弱。

Reasoning-oriented Thinking modes improve most matched model pairs after auditing, but the gains depend on baseline capability and are not uniformly positive. These results suggest that visible deliberation helps only when it remains anchored to units, assumptions, and engineering constraints; otherwise, it may drift from decisive answer boundaries.

经审核发现,面向推理的思维模式(Thinking modes)改善了大多数匹配模型对的表现,但这种提升取决于基准能力,且并非在所有情况下都呈正向。这些结果表明,显式的深思熟虑只有在锚定单位、假设和工程约束时才有效;否则,它可能会偏离确定的答案边界。

WuYuEval therefore provides both an evaluation resource and an empirical basis for developing SWM-oriented foundation models with professional reasoning chains and explicit constraint control.

因此,WuYuEval 不仅提供了一种评估资源,也为开发具备专业推理链和明确约束控制的固体废物管理领域基础模型提供了实证基础。