신뢰성 있는 자율적 설명 가능한 인공지능(XAI)으로 향한 연구: 검증 방법 및 모델의 신뢰성을 평가하기 위한 개방형 벤치마크
Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Better Model Faithfulness
설명 가능한 인공지능(XAI)은 사용자가 모델의 동작을 이해하고 잠재적인 오류를 식별하는 데 도움을 줍니다. 자율적인 XAI 시스템은 대규모 언어 모델(LLM)을 사용하여 자연어 상호작용을 통해 설명을 더욱 쉽게 만들지만, 동시에 그럴듯하지만 실제로는 신뢰성이 떨어지는 설명을 생성할 수도 있습니다. 이는 복잡한 모델에 대한 신뢰할 수 없는 XAI 결과가 LLM에 의해 증폭되어 사용자에게 오해를 불러일으킬 위험이 있기 때문입니다. 본 연구에서는 명시적인 검증을 통해 설명의 신뢰성을 향상시키는 프레임워크인 Faithful Agentic XAI (FAX)를 제안합니다. FAX는 초안 설명을 주장에 분해하고, 내재적으로 신뢰할 수 있는 도구를 사용하여 이러한 주장과 교차 검사를 수행하며, 최종 생성 전에 뒷받침되지 않거나 모순되는 주장을 필터링합니다. 또한, 모델별 신뢰성을 평가하기 위한 복잡한 정책, 다양한 목표 및 어려운 시나리오를 포함하는 개방형 강화 학습 벤치마크인 CRAFTER-XAI-Bench를 소개합니다. CRAFTER-XAI-Bench에서 FAX는 가장 강력한 기준 모델의 시뮬레이션 신뢰성을 0.20에서 0.46으로 향상시키면서 높은 정보성, 관련성 및 유창성을 유지합니다. 세 가지 테이블 기반 벤치마크에서 FAX는 기존의 자율적 XAI 기준 모델과 경쟁력 있는 성능을 보이지만, 분석 결과 이러한 환경은 작업 정확도와 모델별 신뢰성이 혼동될 수 있음을 보여줍니다. 이러한 연구 결과는 명시적인 검증이 신뢰성 있는 자율적 XAI에 필수적이며, 신뢰성 벤치마크는 대상 모델 자체의 동작을 기준으로 설명을 평가하도록 설계되어야 함을 시사합니다.
Explainable AI (XAI) helps users interpret model behavior and identify potential faults. Agentic XAI systems use Large Language Models (LLMs) to make explanations more accessible through natural-language interaction, but they can also produce plausible yet unfaithful explanations. This risk arises because unreliable XAI outputs for complex models can be amplified by LLMs and mislead users. We propose Faithful Agentic XAI (FAX), a framework that improves explanation faithfulness through explicit verification. FAX decomposes draft explanations into claims and cross-checks them against inherently faithful tools, filtering unsupported or contradictory claims before final generation. We also introduce CRAFTER-XAI-Bench, an open-world reinforcement learning benchmark with complex policies, diverse goals, and challenging scenarios for assessing model-specific faithfulness. On CRAFTER-XAI-Bench, FAX improves simulation faithfulness from 0.20 for the strongest baseline to 0.46 while maintaining high informativeness, relevance, and fluency. On three tabular benchmarks, FAX performs competitively with prior Agentic XAI baselines, but our analysis shows that these settings can conflate task accuracy with model-specific faithfulness. These findings show that explicit verification is essential for faithful Agentic XAI and that that faithfulness benchmarks must be designed to test explanations against the behavior of the target model itself.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.