SciHazard: 분해된 위험 점수 기반 과학적 안전 위험 측정 벤치마크
SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring
대규모 언어 모델(LLM)은 과학 연구를 지원하는 데 점점 더 많이 활용되지만, 동시에 위험한 과학 지식을 실제 악용으로 이어질 수 있는 가이드라인으로 변환할 수도 있습니다. 기존의 벤치마크는 종종 현실 세계의 위험과 동떨어진 템플릿 기반 질문을 사용하며, 도메인 지식 없이 LLM을 평가하는 방식을 채택합니다. 이러한 문제를 해결하기 위해, 우리는 실제 환경에 기반한 과학적 위험 벤치마크인 SciHazard와 유해성을 측정하기 위한 데이터셋 독립적인 평가 프레임워크를 소개합니다. SciHazard는 12개 분야에 걸쳐 2400개의 위험 질문과 600개의 안전 관련 질문을 포함하며, 모든 질문은 규제 대상 기관 및 문서화된 실패 사례와 연관되어 있습니다. extsc{DeHarm-Score}를 계산하기 위해, 우리는 질문의 위험 정도, 거부 행동, 그리고 응답 수준의 위험을 결합하는 분해된 평가 절차를 개발했습니다. 거부되지 않은 응답에 대해, 우리는 응답 수준의 유해성을 실행 가능성( extsc{Executability}, 중요도 가중치가 적용된 동적 체크리스트로 측정)과 새로운 위험 요소( extsc{Net-new risk}, 검색 증강 기반 주장 추출 및 합성 장벽 검증을 통해 평가)로 더욱 세분화합니다. 전문가 검증 연구 결과, extsc{DeHarm-Score}는 가장 강력한 기준 모델보다 90.17% 더 높은 수준의 일치도를 보였습니다. 우리는 31개의 최첨단 LLM과 연구 에이전트를 광범위한 과학적 안전 평가에 사용했습니다. 주목할 점은, 연구 에이전트는 표준 LLM보다 평균 extsc{DeHarm-Score}가 32.3% 더 높았으며, 이는 자율 에이전트가 현재의 안전 방어 체계에서 간과되고 있는 중요한 영역임을 보여줍니다. 코드 및 데이터셋은 https://anonymous.4open.science/r/DeharmScore-7B55 에서 제공됩니다.
Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance. Existing benchmarks often rely on templated queries disconnected from real-world hazards, and employ LLM-as-a-Judge paradigms without domain grounding. To address this, we introduce SciHazard, a real-world-grounded benchmark for scientific risks and a dataset agnostic evaluation framework for measuring harmfulness. SciHazard contains 2400 hazardous questions and 600 oversafety questions across 12 disciplines, with both queries grounded in regulated entities and documented failure scenarios. To compute \textsc{DeHarm-Score} , we develop a decomposed evaluating procedure that combines query hazard severity, refusal behavior, and response-level risk. For non-refused responses, it further decomposes response-level harm into \textsc{Executability}, quantified via dynamic checklists with importance weighting, and \textsc{Net-new risk}, assessed through retrieval-augmented claim extraction and synthesis-barrier verification. An expert-validation study shows that \textsc{DeHarm-Score} improves agreement with expert annotations by 90.17\% over the strongest baseline. We benchmark 31 frontier LLMs and deep research agents in an extensive scientific safety evaluation. Notably, deep research agents yield 32.3\% higher mean \textsc{DeHarm-Score} than standard LLMs, exposing autonomous agents as a critical blind spot in current safety defenses. Code and dataset are available at https://anonymous.4open.science/r/DeharmScore-7B55.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.