2607.17745v1 Jul 20, 2026 cs.AI

WuYu-EnvLE-Bench: 환경 법 집행 분야의 대규모 언어 모델 평가를 위한 벤치마크

WuYu-EnvLE-Bench: A Benchmark for Evaluating Large Language Models in Environmental Law Enforcement

Yi Zhang
Yi Zhang
Citations: 8
h-index: 1
Ziliang Yang
Ziliang Yang
Citations: 0
h-index: 0
Jiachao Ke
Jiachao Ke
Citations: 84
h-index: 4
Zongguo Wen
Zongguo Wen
Citations: 395
h-index: 10
K. Lin
K. Lin
Citations: 0
h-index: 0
Haihong Xu
Haihong Xu
Citations: 0
h-index: 0

대규모 언어 모델(LLM)은 점점 더 환경 단속에 활용되고 있지만, 이러한 모델이 추적 가능한 단속 결정을 내릴 수 있는 능력은 아직 명확하지 않습니다. 본 연구에서는 실제 단속 사례, 규제 표준 및 전문가 검토를 기반으로 구축된 WuYu-EnvLE-Bench 벤치마크를 소개합니다. 이 벤치마크는 사전 단속, 단속 과정, 사후 단속 워크플로우 전반에 걸쳐 2,521개의 테스트 케이스, 14가지 작업 및 12가지 오염 매개체 하위 영역을 포함합니다. 절대적인 환경 단속 점수(AES)와 지능형 단속 지수(IEI)를 사용하여 오픈 소스 및 폐쇄 소스 LLM의 기능, 응답 품질 및 리소스 효율성을 평가했습니다. 결과는 LLM이 규칙 기반 작업에서는 잘 수행되지만, 증거 연결 구축, 모순 감지, 다중 출처 통합 및 절차적 판단에서는 신뢰성이 떨어진다는 것을 보여줍니다. 모델 크기 확장도 감소하는 수익률을 보입니다. 중간 규모의 모델은 구조화된 작업에서 선도적인 모델에 근접하는 반면, 더 큰 모델은 증거 기반 추론의 한계를 안정적으로 극복하지 못합니다. WuYu-EnvLE-Bench는 증거 기반, 규칙 인식 및 작업 적응형 단속 추론의 필요성을 강조합니다.

Original Abstract

Large language models (LLMs) are increasingly considered for environmental enforcement, but their ability to produce traceable enforcement decisions remains unclear. We introduce WuYu-EnvLE-Bench, a benchmark built from real enforcement cases, regulatory standards, and expert review. It contains 2,521 benchmark instances, 14 tasks, and 12 pollution-medium subdomains across pre-enforcement, in-enforcement, and post-enforcement workflows. Using Absolute Environmental Enforcement Score (AES) and Intelligent Enforcement Index (IEI), we evaluate open-source and closed-source LLMs across capability, response quality, and resource efficiency. Results show that LLMs perform well on rule-bounded tasks but remain unreliable in evidence-chain construction, contradiction detection, multi-source integration, and procedural judgment. Model scaling also shows diminishing returns: medium-sized models approach leading models in structured tasks, while larger models do not reliably overcome evidence-reasoning bottlenecks. WuYu-EnvLE-Bench highlights the need for evidence-grounded, rule-aware, and task-adaptive enforcement reasoning.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!