특징에서 행동으로: 전통적인 AI 시스템과 자율형 AI 시스템에서의 설명 가능성
From Features to Actions: Explainability in Traditional and Agentic AI Systems
지난 10년 동안 설명 가능한 AI는 주로 개별 모델 예측을 해석하는 데 초점을 맞추어 왔으며, 고정된 의사 결정 구조 하에서 입력과 출력을 연결하는 사후 설명을 생성합니다. 최근 대규모 언어 모델(LLM)의 발전으로 인해 다단계 경로를 통해 행동이 전개되는 자율형 AI 시스템이 등장했습니다. 이러한 환경에서 성공과 실패는 단일 출력보다는 의사 결정의 연속에 의해 결정됩니다. 기존의 설명 방법이 정적인 예측에 적합하지만, 시간이 지남에 따라 행동이 나타나는 자율형 환경에서는 어떻게 적용될 수 있는지 불분명합니다. 본 연구에서는 속성 기반 설명과 추적 기반 진단을 비교하여 정적 및 자율형 환경 모두에서 설명 가능성 간의 간극을 메우고자 합니다. 이 차이점을 명확히 하기 위해, 본 연구에서는 정적 분류 작업에 사용되는 속성 기반 설명과 자율형 벤치마크(TAU-bench Airline 및 AssistantBench)에 사용되는 추적 기반 진단을 경험적으로 비교합니다. 연구 결과, 속성 방법은 정적 환경에서 안정적인 특징 순위를 달성하는 반면(Spearman $ρ= 0.86$), 자율형 경로에서의 실행 수준 실패를 진단하는 데 신뢰성 있게 적용될 수 없습니다. 반면, 자율형 환경에 대한 추적 기반 루브릭 평가는 행동 오류를 지속적으로 특정하고, 실패한 실행에서 상태 추적 불일치가 2.7배 더 흔하며, 성공 확률을 49% 감소시킨다는 것을 보여줍니다. 이러한 결과는 자율형 시스템의 자율적인 AI 행동을 평가하고 진단할 때 경로 수준의 설명 가능성으로의 전환을 촉구합니다.
Over the last decade, explainable AI has primarily focused on interpreting individual model predictions, producing post-hoc explanations that relate inputs to outputs under a fixed decision structure. Recent advances in large language models (LLMs) have enabled agentic AI systems whose behaviour unfolds over multi-step trajectories. In these settings, success and failure are determined by sequences of decisions rather than a single output. While useful, it remains unclear how explanation approaches designed for static predictions translate to agentic settings where behaviour emerges over time. In this work, we bridge the gap between static and agentic explainability by comparing attribution-based explanations with trace-based diagnostics across both settings. To make this distinction explicit, we empirically compare attribution-based explanations used in static classification tasks with trace-based diagnostics used in agentic benchmarks (TAU-bench Airline and AssistantBench). Our results show that while attribution methods achieve stable feature rankings in static settings (Spearman $ρ= 0.86$), they cannot be applied reliably to diagnose execution-level failures in agentic trajectories. In contrast, trace-grounded rubric evaluation for agentic settings consistently localizes behaviour breakdowns and reveals that state tracking inconsistency is 2.7$\times$ more prevalent in failed runs and reduces success probability by 49\%. These findings motivate a shift towards trajectory-level explainability for agentic systems when evaluating and diagnosing autonomous AI behaviour. Resources: https://github.com/VectorInstitute/unified-xai-evaluation-framework https://vectorinstitute.github.io/unified-xai-evaluation-framework
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.