과거 지향성: 왜 GUI 기반 모델이 과거에 갇히는가?
Prior Directions: Why GUI Grounding Gets Locked in the Past
시각-언어 모델은 종종 현재 장면의 결정을 내리기 위해 이전 시각적 상태에 대한 설명을 사용합니다. 장면이 변경되면 오래된 언어 정보가 올바른 시각적 판단을 잘못된 답변으로 이끌 수 있습니다. 본 연구에서는 이러한 실패를 '시각적 고착 현상'이라고 정의하고, 제어된 환경에서 오직 언어적인 이전 정보만 변화시키는 방식으로 이를 분석합니다. 다양한 모델에서, 더 강한 고착 현상은 모델 표현이 최종 답변 전에 작게 변할 때 나타납니다. 이러한 반전은 고착 현상이 해당 표현의 이동 거리 자체가 아니라 그 이동 방식에 달려 있음을 시사합니다. 수정하기 어려운 모델에서는 이전 정보로 인한 변화가 반복적으로 나타나는 일련의 방향, 즉 'Prior Directions(과거 지향성)'를 따라 집중됩니다. 이러한 과거 지향성은 테스트 데이터에서도 반복적으로 나타나며, 4개의 모델을 비교한 결과는 더 높은 집중도가 더 강한 고착 현상과 관련이 있음을 보여줍니다. 제어된 실험 결과, 과거 지향성과 일치하는 요소를 제거하면 시각적 기반 연결(grounding)이 복원되는 반면, 동일한 크기의 수직 성분을 제거해도 큰 영향은 없습니다. 따라서 이전 정보의 통제는 이전 정보로 인한 변화가 답변을 생성하는 데 사용되는 표현에서 응집되고 재사용 가능한 패턴을 형성할 때 발생합니다. 이러한 설명은 동일한 이전 정보가 어떤 모델에서는 수정 가능하지만 다른 모델에서는 지배적인 요소가 되는 이유를 설명합니다.
Vision-language models often use descriptions of earlier visual states to make decisions about the current scene. When the scene changes, stale language can redirect an otherwise correct visual judgment toward an outdated answer. We study this failure as visual lock-in in a controlled grounding setting where only the verbalized prior varies. Across models, stronger lock-in accompanies smaller changes in the model representation before the final answer. This reversal suggests that lock-in depends not on how far this representation moves, but on how that movement is organized. In models that are harder to correct, prior-induced changes concentrate along a compact set of directions that repeatedly appear across examples. We call these recurrent axes the Prior Directions. They recur on held-out examples, while a descriptive four-model comparison associates greater concentration with stronger lock-in. Controlled interventions show that removing the component aligned with the Prior Directions restores visual grounding, whereas removing an equally large orthogonal component has little effect. Prior control thus arises when prior-induced changes form a coherent and reusable pattern in the representation used to produce the answer. This account explains why the same prior remains revisable in one model yet becomes dominant in another.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.