DiagFlowBench: 언어 모델이 근거 기반 진단 대화에서 절차를 벗어난 입력에 어떻게 대응하는지 평가
DiagFlowBench: Evaluating How Language Models Handle Off-Procedure Inputs in Grounded Diagnostic Dialogue
언어 모델은 점점 더 많은 분야에서 자문 시스템으로 활용되고 있으며, 특히 유지 보수 작업에서 중요한 역할을 수행합니다. 환각 현상을 방지하기 위해 최근에는 이러한 모델을 절차 문서와 연결하여 승인된 단계 내에서만 작동하도록 제한하는 경향이 있습니다. 그러나 실제 운영 환경에서는 작업자의 질문이 종종 이러한 경로를 벗어나는 경우가 많으며, 이 경우 모델은 대화 중에 범위를 벗어난 입력을 인식해야 합니다. 현재의 벤치마크는 이러한 역동성을 충분히 고려하지 못합니다. 본 연구에서는 소비자 제조 업체의 산업용 진단 흐름도를 기반으로 구성된 50개의 흐름도와 이를 활용하여 생성된 1,676개의 다중 회전 대화 데이터셋인 DiagFlowBench를 소개합니다. 이 데이터셋은 절차를 준수하는 발언과 범위를 벗어난 발언을 비교합니다. 십 개 이상의 상용 및 오픈 소스 모델에 대한 평가 결과, 모델별로 상당한 차이를 보이는 거부율을 확인했습니다. 많은 모델이 사실이지만 문맥상 부적절한 단계를 선택하는 경향이 있으며, 이는 허구의 정보를 생성하는 것보다 발생합니다. 이러한 매핑된 정보는 타당해 보이고 권위적으로 들릴 수 있지만, 실제로는 잘못된 조언이며, 이는 근거 시스템에 대한 심각한 취약점을 드러냅니다.
Language models increasingly serve as advisory systems in maintenance operations. To prevent hallucination, recent systems ground these models in procedural documentation to constrain them to approved steps. In practice, however, operator queries frequently stray from this path, requiring models to recognise out-of-scope inputs mid-conversation, a dynamic that current benchmarks rarely prioritise. We introduce DiagFlowBench, a dataset of 50 industrial diagnostic flowcharts from a consumer manufacturer converted into 1,676 multi-turn conversations that contrast compliant with out-of-scope utterances. Evaluating a panel of ten commercial and open-weight models reveals high variability in abstention rates, with models commonly selecting a real but contextually inadequate step rather than fabricating facts. The inherent plausibility and authority of this mapped but wrong advice exposes a challenging vulnerability for grounding systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.