OpenSafeIntent: 이중 용도 프롬프트 집합에서 의도 기반 안전 완성을 평가
OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets
안전한 모델 응답은 유용한 도움을 제공하는 동시에 해를 끼치지 않아야 하지만, 독립적인 프롬프트를 사용한 평가로는 이러한 특성을 파악하기 어렵습니다. 본 논문에서는 의도를 다양하게 변화시키면서 기본 작업은 고정한 통제된 프롬프트 집합으로 구성된 벤치마크인 OpenSafeIntent를 소개합니다. 각 데이터 포인트는 동일한 작업의 양성, 이중 용도, 악의적 변형을 포함합니다. 이러한 설계 덕분에 모델이 평균적으로 안전해 보이는지 평가하는 것이 아니라, 의도 변화에 따른 응답의 적절성을 평가할 수 있습니다. 다양한 모델들을 실험한 결과, 프롬프트 수준의 안전성은 중요한 실패를 숨길 수 있음을 확인했습니다. 모델들은 종종 일관된 의도를 가진 변형된 프롬프트에서도 안전하지 않은 결과를 생성하며, 이중 용도의 특성이 문장 재구성(paraphrase)에 취약하고, 위험한 주제에 대한 고수준 답변이 항상 안전하지 않으며, 모호한 요청을 더 안전한 작업으로 재구성하는 응답은 안전 경계를 위반할 가능성이 훨씬 낮다는 것을 발견했습니다. 이러한 결과는 안전한 모델 응답이 독립적인 프롬프트에 대한 단일한 안전성-유용성 균형으로 평가되는 것이 아니라, 통제된 작업 변형을 통해 의도 기반의 적절성을 평가해야 함을 시사합니다.
Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts. We introduce OpenSafeIntent, a benchmark of controlled prompt-sets that vary intent while holding the underlying task fixed. Each datapoint contains benign, dual-use, and malicious variants of the same task. This design lets us evaluate whether models calibrate assistance across intent shifts, rather than merely appearing safe on average. Across a broad model suite, we find that prompt-level safety hides important failures: models often fail to remain safe across matched intent variants, dual-use behavior is brittle under paraphrase, high-level answers on risky topics are not reliably safe, and responses that reframe ambiguous requests into safer tasks are substantially less likely to cross the safety boundary. Our results suggest that safe completion should be evaluated as intent-calibrated behavior over controlled task variants, not as a single safety-helpfulness tradeoff over independent prompts.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.