SoCRATES: 도메인 및 사회인지적 변형을 고려한 능동적인 LLM 중재의 신뢰성 있는 자동 평가를 향하여
SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations
LLM 중재자의 평가는 여전히 어려운 과제입니다. 중재는 분쟁 당사자들의 변화하는 감정, 의도 및 맥락에 의해 실시간으로 형성되는 과정이기 때문입니다. 기존 테스트 환경은 소수의 전문가가 작성한 도메인에 의존하며, 주로 전략적 태도를 변형시키고 각 발언을 모든 주제에 대해 평가하여 관련 없는 잡음을 유발합니다. 본 연구에서는 실제 다중 도메인 환경에서 능동적인 LLM 중재자를 평가하기 위한 벤치마크인 SoCRATES를 소개합니다. SoCRATES는 에이전트 기반 파이프라인을 통해 실제 갈등으로부터 시나리오를 구성하고, 전략적 태도, 당사자 구성, 이력 길이, 감정 반응성 및 문화적 정체성을 포함한 다섯 가지 사회인지적 적응 축을 탐구하며, 각 주제에 대해 해당 주제를 발전시키는 발언만을 주제별 평가기로 평가합니다. 제안하는 평가기는 인간 전문가와의 일치도가 0.82로 나타났으며, 이는 턴 단위 기준으로 측정된 기존 방법보다 두 배 이상 높은 수치입니다. SoCRATES를 사용하여 최첨단 LLM 8개를 비교평가한 결과, 가장 우수한 중재자조차도 다양한 실제 환경에서 약 3분의 1 정도의 합의 격차만을 해소할 수 있다는 사실을 확인했습니다. 또한 성능은 사회인지적 축에 따라 크게 달라졌으며, 이는 다양한 조건에 대한 사회적 적응이 향후 발전의 중요한 방향임을 시사합니다.
Evaluating LLM mediators remains challenging, as mediation unfolds as a real-time trajectory shaped by disputants' shifting emotions, intentions, and context. Existing testbeds rely on a few expert-authored domains, vary mainly strategic posture, and score every turn against every topic, introducing off-topic noise. We introduce SoCRATES, a benchmark for evaluating proactive LLM mediators in realistic, multi-domain testbeds. It constructs scenarios from real conflicts through an agentic pipeline across eight domains, probes five socio-cognitive adaptation axes (strategic posture, party composition, history length, emotional reactivity, and cultural identity), and scores each topic only on the turns that advance it via a topic-localized evaluator. The evaluator reaches 0.82 alignment with human experts, more than doubling a per-turn baseline. Benchmarking eight frontier LLMs, we find that even the strongest mediator closes only about a third of the unmediated consensus gap under diverse and realistic testbeds, with performance varying sharply by socio-cognitive axis, highlighting that progress lies in social adaptation to diverse conditions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.