StepGuard: 단일 단계 보정 기반 웹 탐색 보호
StepGuard: Guarding Web Navigation via Single-Step Calibration
웹 탐색은 에이전트가 자연어 목표를 따르고, 웹 페이지와 상호 작용하며, 정확한 답변을 생성하도록 요구합니다. 최근 연구에서는 시각-언어 모델과 강화 학습 기술이 활용되었지만, 기존 방법들은 여전히 보상 불일치 및 오류 전파로 인한 단일 단계 취약성을 가지고 있습니다. 이러한 보상 문제를 해결하기 위해, 우리는 탐색 우선 모드와 질문 답변 우선 모드를 동적으로 전환하여 보상 충돌을 완화하는 Dynamic Dual-Policy Optimization (DDPO)를 설계했습니다. 또한, 단일 단계 오류를 교정하기 위해, 각 단계별 신뢰도를 추정하고, 필요한 경우에만 반성을 트리거하며, 대조 학습 기반의 보상을 사용하여 자체 수정 및 정확도 향상을 유도하는 Confidence-Guided Adaptive Navigation Reflection (CANR) 메커니즘을 제안합니다. 위에서 설명한 주요 구성 요소를 바탕으로, 우리는 단일 단계 보정을 통한 웹 탐색 보호 프레임워크인 StepGuard를 개발했습니다. 실험 결과, 우리 방법은 탐색 및 답변 정확도를 크게 향상시켜 표준 웹 탐색 벤치마크에서 새로운 최고 성능을 달성함을 보여줍니다.
Web navigation requires agents to follow natural language goals, interact with web pages, and produce accurate answers. While recent advances leverage vision-language models and reinforcement learning, existing methods still suffer from single-step fragility due to reward misalignment and error propagation. To tackle the reward entanglement, we design Dynamic Dual-Policy Optimization (DDPO), which dynamically switches between a navigation-first mode for exploration and an answer-first mode for question-answering to mitigate reward conflict. To calibrate the single-step error, we propose Confidence-Guided Adaptive Navigation Reflection (CANR), a mechanism that estimates per-step confidence, triggers reflection only when necessary, and uses contrastive rewards to encourage self-correction to calibrate the single-step inaccuracy. With the above as the main components, we finally develop our StepGuard, a new framework of Guarding Web Navigation via Single-Step Calibration. Experiments demonstrate that our approach significantly improves navigation and answer accuracy, setting new state-of-the-art performance on standard web navigation benchmarks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.