그리드 진단을 위한 다중 모드 대규모 언어 모델(LLM)의 작업 조건부 충실성 감사
Task-Conditional Faithfulness Auditing of Multimodal LLMs for Grid Diagnosis
다중 모드 대규모 언어 모델(LLMs)은 토폴로지, 측정값 및 사고 관련 텍스트를 결합하여 그리드 진단에 활용될 수 있지만, 답변 정확도가 작업에 적합한 증거가 사용되었음을 보장하지는 않습니다. 본 논문에서는 작업 조건부 충실성 감사를 수행하기 위한 일반적인 프레임워크를 제안합니다. 이 프레임워크는 자기 보고된 의존도, 개입을 통해 얻은 행동적 의존도 및 사전에 등록된 공학적 중요도를 비교합니다. 먼저 작업별 증거 요구 사항을 등록하고 이를 자기 보고된 의존도와 통제된 모드 제거 조건 하에서의 행동 변화와 비교합니다. 감지된 불일치를 해결하기 위해, 우리는 실패한 응답을 증거 제약 조건 하에서 재생성하고 독립적으로 다시 모드를 제거하여 성능 저하 없이 개선된 근거를 검증하는 증거 기반 수정 및 재감사 메커니즘을 설계했습니다. 사례 연구에서는 IEEE 39 및 118 버스 시나리오에서 서로 다른 규모의 세 가지 LLM을 평가합니다. 이러한 결과는 프레임워크가 작업 조건부 충실성 실패를 감지, 진단 및 수정하는 능력을 검증합니다.
Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish that task-appropriate evidence was used. This letter proposes a general framework in order to conduct task-conditional faithfulness audit. It compares self-reported reliance, intervention-derived behavioral reliance, and preregistered engineering importance. The framework first registers task-specific evidence requirements and compares them with self-reported reliance and behavioral changes under controlled modality ablations. To resolve detected discrepancies, we design an evidence-gated correction and re-audit mechanism that regenerates failed responses under evidence constraints and independently re-ablates them to verify improved grounding without performance loss. Case studies evaluate three differently scaled LLMs on IEEE 39- and 118-bus scenarios. These results validate the framework ability to detect, diagnose, and correct task-conditional faithfulness failures.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.