2607.12640v1 Jul 14, 2026 cs.AI

작은 언어 및 시각-언어 모델 웹 에이전트에서 GRPO의 학습률 제한적 실패: 통제된 가설과 그 메커니즘

A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism

Zhixi Cai
Zhixi Cai
Citations: 87
h-index: 6
Chengguang Gan
Chengguang Gan
Citations: 194
h-index: 7
Yunhao Liang
Yunhao Liang
Citations: 1
h-index: 1
Hanjun Wei
Hanjun Wei
Citations: 264
h-index: 9
Shiwen Ni
Shiwen Ni
Citations: 15
h-index: 2
Qinghao Zhang
Qinghao Zhang
Citations: 115
h-index: 4

검증 가능한 보상을 사용하는 강화 학습, 특히 그룹 상대 정책 최적화(GRPO)는 현재 지도 학습 체크포인트에서 강력한 에이전트를 생성하기 위해 널리 사용됩니다. 본 연구에서는 40억에서 80억 규모의 작은 언어 및 시각-언어 모델 웹 에이전트에 GRPO가 실제로 어떤 영향을 미치는지, 또는 기존 지도 학습 모델의 행동을 단순히 재구성하는지에 대해 질문합니다. 학습률, KL 가중치, 초기화, 클리핑 등 다양한 요소를 변경한 18개의 실험 설정을 통해, 대부분의 작업에서 이미 높은 성능을 보이는 강력한 지도 학습 기준 모델의 성공률을 GRPO가 유의미하게 향상시키지 못한다는 것을 확인했습니다. 특히, 중간에서 높은 학습률에서는 오히려 성능이 저하되는 경향이 나타났습니다. 이러한 결과는 쌍대 비교 테스트, 25개의 평가 시드, 6개의 학습 시드, 레시피 변경, 텍스트 및 Set-of-Marks 스크린샷 관찰을 통해 검증되었으며, 모델 크기를 80억으로 확장해도 동일한 경향이 나타났습니다. 유의미한 성능 저하는 텍스트 데이터에서 주로 관찰되며, Set-of-Marks에서는 미미한 수준입니다. 파이프라인 오류 여부를 확인하기 위해, 동일한 환경, 보상 및 레시피를 사용하여 샘플링을 통해 보상을 얻을 수 있는 작업에 대해 실험한 결과, 성공률이 22% 포인트 증가하는 것을 확인했습니다 (0을 배제하는 신뢰 구간). 따라서 GRPO는 정책의 성능이 탐욕적인 방식보다 더 높은 잠재력을 가진 경우에만 효과적입니다. 본 연구에서는 이러한 실패 원인을 분석합니다. 중간 학습률은 에이전트의 성능을 저하시키고, 높은 학습률은 에이전트를 완전히 붕괴시킵니다. 두 가지 현상은 서로 다른 방식으로 나타나는데, 낮은 성능을 보이는 경향은 주로 어텐션 및 MLP 블록과 관련되어 있으며, 붕괴 현상은 특정 그룹으로 추적할 수 없습니다. 또한, 가중치 변화에서 가장 큰 영향을 미치는 임베딩 변경은 인과 관계가 없는 것으로 나타났습니다. 40억 규모의 모델에서는 후반 레이어에서의 유효 순위가 성능을 반영하지만, 80억 규모의 모델에서는 두 요소가 분리되는 경향이 있습니다. 이러한 상관관계는 작은 모델에만 해당되므로, 이는 크기에 따라 달라지는 현상으로 보고합니다.

Original Abstract

Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent. We ask whether it adds skill to a small language and vision-language model web agent at the 4B to 8B scale, or whether it mostly reshapes behavior the supervised model already has. Across a control grid of 18 runs that varies learning rate, KL weight, seed, initialization, and clipping, no configuration credibly improves the success rate of a strong supervised baseline on tasks the agent has largely mastered. On the text track, moderate to high learning rates make it credibly worse. The null holds under paired testing, 25 evaluation seeds, 6 training seeds, changes to the recipe, both text and Set-of-Marks screenshot observations, and scaling the backbone to 8B; the credible harm is a text-track finding and is only nominal under Set-of-Marks. To show that the null reflects the setting and not a broken pipeline, we run the identical harness, reward, and recipe on tasks whose reward is reachable by sampling, and there the success rate rises by 22 points with a paired interval that excludes zero. GRPO therefore helps only when there is headroom to climb, meaning the sampled policy already succeeds more often than the greedy one. We then explain the failure. A middle learning rate degrades the agent and a high one collapses it, and the two regimes form a double dissociation: grafting localizes the degrade regime to the attention and MLP blocks, while the collapse regime cannot be traced to any single group, and the embedding change that dominates the weight movement is causally inert. At 4B, effective rank in the late layers tracks capability in both directions; at 8B the two come apart. This coupling is specific to the smaller model, so we report it as scale-dependent.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!