Janus: 목표 지향적 정보 왜곡에 대한 LLM 성능 평가 기준
Janus: A Benchmark for Goal-Conditioned Information Distortion in LLMs
LLM의 기만성은 종종 조작된 주장, 명백한 거짓말 또는 전략적인 은폐와 같은 직접적인 지표를 통해 평가됩니다. 그러나 많은 실제 오해를 불러일으키는 의사소통은 허위 진술에 의존하지 않고, 오히려 사실 정보를 선택적으로 다루는 방식으로 발생합니다. 여기에는 불리한 증거 누락, 불리한 세부 사항 완화, 유리한 세부 사항 강조 또는 정확한 자격 요건을 모호한 표현으로 대체하는 것이 포함됩니다. 기존의 평가 기준은 이러한 미묘하고 잠재적으로 더 위험한 실패 양상을 대부분 간과합니다. 본 논문에서는 사실에 기반한 LLM 출력에서 목표 지향적인 담론 왜곡을 측정하기 위한 평가 기준인 JANUS를 소개합니다. 저희 평가 기준의 각 시나리오는 유리하고 불리한 사실 풀을 고정된 상태로 제공하며, 중립 조건과 목표 지향적 조건을 비교합니다. 목표는 채택, 등록, 승인 또는 지지를 증가시키는 것이지만, 이는 직접적으로 영향을 받는 개인이나 집단에게 잠재적인 피해를 초래할 수 있습니다. JANUS는 모든 출력이 동일한 사실 풀을 사용하도록 제한하여, 환각 및 조작으로부터 오해의 소지가 있는 전체적인 인상을 분리합니다. JANUS는 8개 도메인에 걸쳐 총 160개의 시나리오로 구성되어 있으며, 각 시나리오에는 중립적 프롬프트와 목표 지향적 프롬프트가 함께 제공되며, 관련된 사실 정보가 주석으로 첨부되어 있습니다. 12개의 LLM을 대상으로 실시한 광범위한 실험 결과, 일관된 목표 지향적인 왜곡이 나타났습니다. 이는 현재 모델이 여전히 인센티브 및 프레임 지향적인 목표에 민감하며, 선택적으로 오해를 불러일으키는 의사소통에 대한 강력한 안전장치가 부족하다는 것을 보여줍니다. 저희는 향후 연구를 위해 데이터셋과 코드를 공개합니다.
LLM deception is often evaluated through direct markers such as fabricated claims, explicit lies, or strategic concealment. However, many real-world misleading communications do not depend on false statements, rather, they arise from selective treatment of true material facts: omitting adverse evidence, softening unfavorable details, emphasizing favorable details, or replacing precise qualifications with vague language. Existing benchmarks largely miss this subtler and arguably more dangerous failure mode. We introduce JANUS, a benchmark for measuring goal-conditioned pragmatic distortion in fact-grounded LLM outputs. Each scenario in our benchmark provides a fixed pool of favorable and adverse facts and compares a neutral condition against a goal-directed condition, such as increasing adoption, enrollment, approval, or support, despite potential harm to directly affected individuals or groups. Because all outputs are constrained to use the same fact pool, JANUS isolates misleading net impressions from hallucination and fabrication. JANUS contains 160 scenarios across 8 domains, with each scenario paired with neutral and goal-conditioned prompts and annotated material facts. Extensive experiments across 12 LLMs reveal consistent goal-conditioned distortions, demonstrating that current models remain sensitive to incentive and framing objectives and lack robust safeguards against selectively misleading communication. We publicly release our corpus and code for future research.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.