WuYu-EnvLE-Bench: 환경 법 집행 분야의 대규모 언어 모델 평가를 위한 벤치마크
WuYu-EnvLE-Bench: A Benchmark for Evaluating Large Language Models in Environmental Law Enforcement
대규모 언어 모델(LLM)은 점점 더 환경 단속에 활용되고 있지만, 이러한 모델이 추적 가능한 단속 결정을 내릴 수 있는 능력은 아직 명확하지 않습니다. 본 연구에서는 실제 단속 사례, 규제 표준 및 전문가 검토를 기반으로 구축된 WuYu-EnvLE-Bench 벤치마크를 소개합니다. 이 벤치마크는 사전 단속, 단속 과정, 사후 단속 워크플로우 전반에 걸쳐 2,521개의 테스트 케이스, 14가지 작업 및 12가지 오염 매개체 하위 영역을 포함합니다. 절대적인 환경 단속 점수(AES)와 지능형 단속 지수(IEI)를 사용하여 오픈 소스 및 폐쇄 소스 LLM의 기능, 응답 품질 및 리소스 효율성을 평가했습니다. 결과는 LLM이 규칙 기반 작업에서는 잘 수행되지만, 증거 연결 구축, 모순 감지, 다중 출처 통합 및 절차적 판단에서는 신뢰성이 떨어진다는 것을 보여줍니다. 모델 크기 확장도 감소하는 수익률을 보입니다. 중간 규모의 모델은 구조화된 작업에서 선도적인 모델에 근접하는 반면, 더 큰 모델은 증거 기반 추론의 한계를 안정적으로 극복하지 못합니다. WuYu-EnvLE-Bench는 증거 기반, 규칙 인식 및 작업 적응형 단속 추론의 필요성을 강조합니다.
Large language models (LLMs) are increasingly considered for environmental enforcement, but their ability to produce traceable enforcement decisions remains unclear. We introduce WuYu-EnvLE-Bench, a benchmark built from real enforcement cases, regulatory standards, and expert review. It contains 2,521 benchmark instances, 14 tasks, and 12 pollution-medium subdomains across pre-enforcement, in-enforcement, and post-enforcement workflows. Using Absolute Environmental Enforcement Score (AES) and Intelligent Enforcement Index (IEI), we evaluate open-source and closed-source LLMs across capability, response quality, and resource efficiency. Results show that LLMs perform well on rule-bounded tasks but remain unreliable in evidence-chain construction, contradiction detection, multi-source integration, and procedural judgment. Model scaling also shows diminishing returns: medium-sized models approach leading models in structured tasks, while larger models do not reliably overcome evidence-reasoning bottlenecks. WuYu-EnvLE-Bench highlights the need for evidence-grounded, rule-aware, and task-adaptive enforcement reasoning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.