읽기는 추론이 아니다: 시각-텍스트 압축에서의 주체적 정책 격차 해소
Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression
다단계 언어 모델 에이전트는 방대한 상호 작용 기록을 반복적으로 처리하므로 상당한 컨텍스트 비용이 발생합니다. 시각-텍스트 압축은 이 기록을 이미지로 변환하여 이러한 비용을 줄이지만, 결과적인 모달리티 변화는 뚜렷한 능력 격차를 야기합니다. 역사 정보 복구, 동일 상태에서의 의사 결정, 전체 경로에 대한 통제된 평가를 통해, 우리는 이러한 격차가 OCR 품질만으로는 설명될 수 없음을 보여줍니다. 시각-역사 에이전트는 행동 선택, 질의 구성, 중단, 증거 활용에서 체계적인 편향을 보이는 경향이 있으며, 이는 주체적 정책 격차를 드러냅니다. 본 연구에서는 extbf{CAPS (C}ross-modal extbf{A}gentic extbf{P}olicy extbf{S}elf-distillation)라는 두 단계의 교차 모달 에이전트 정책 자체 증류 프레임워크를 제안합니다. 이 프레임워크는 동일 모델의 더 강력한 텍스트-역사 정책을 사용하여 시각-역사 에이전트를 지도 학습합니다. 오프라인 경로 자체 증류는 성공적인 텍스트 정책 행동을 시각-역사 입력에 전달하고, 온라인 정책 자체 증류는 강화 학습 과정에서 시각-역사 정책이 방문하는 상태에 대한 밀집적인 지도를 제공합니다. SearchQA 데이터셋에서 CAPS는 3B 및 7B 모델을 사용할 때 각각 5.0%와 3.4%의 성능 향상을 보입니다. 전체 역사 ALFWorld 데이터셋에서는 해당 성능 향상이 각각 15.6%와 14.5%입니다. 다양한 환경에서 CAPS는 평균 메모리-컨텍스트 비용을 최대 63.3%, 최고 비용을 최대 83.4%까지 줄였습니다 (동일한 텍스트-역사 정책과 비교). 이러한 결과는 명시적인 교차 모달 정책 자체 증류가 시각-텍스트 압축 하에서 에이전트의 능력을 유지할 수 있음을 보여줍니다. 본 연구의 코드는 추후 공개될 예정입니다.
Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduces these costs by rendering history as images, but the resulting modality shift creates a marked capability gap. Through controlled evaluations of history recovery, matched-state decisions, and complete trajectories, we show that this gap cannot be explained by OCR quality alone. Visual-history agents exhibit systematic drift in action selection, query formulation, stopping, and evidence use, revealing an agentic policy gap. We introduce \textbf{CAPS}, a two-stage \textbf{C}ross-modal \textbf{A}gentic \textbf{P}olicy \textbf{S}elf-distillation framework that uses the same model's stronger text-history policy to supervise its visual-history counterpart. Offline trajectory self-distillation transfers successful text-policy behavior to visual-history inputs, while online policy self-distillation provides dense supervision on states visited by the visual-history policy during reinforcement learning. On SearchQA, CAPS improves over AgentOCR by 5.0\% and 3.4\% with 3B and 7B backbones, respectively. On full-history ALFWorld, the corresponding gains are 15.6\% and 14.5\%. Across settings, CAPS reduces average memory-context cost by up to 63.3\% and peak cost by up to 83.4\% relative to matched text-history policies. These results show that explicit cross-modal policy self-distillation can preserve agent capability under vision--text compression. Our code will be made publicly available in a future release.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.