라우팅은 가장 가치 있는 곳에서 가장 학습하기 어렵다: 웹 에이전트를 위한 표현 라우팅의 경계
Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents
웹 에이전트는 브라우저를 텍스트, 픽셀 또는 둘 다를 통해 관찰하며, 이러한 관찰 방식은 일반적으로 모든 작업에 대해 한 번 선택됩니다. 본 연구에서는 VisualWebArena 및 WebArena에서 8개의 사이트-모델 조합(셀)에 걸쳐 6가지 관찰 방식을 측정하고, 각 작업을 위해 관찰 방식을 개별적으로 선택했을 때 얻을 수 있는 이점을 분석했습니다. 이러한 방식들은 상호 보완적입니다. 즉, 다른 방식으로는 해결할 수 없는 문제를 해결하며, 구조적으로 다른 방식으로 실패하고, 최적의 선택은 작업 세트에 따라 달라집니다. 명백한 장점은 모든 작업에 대해 우수한 관찰 방식을 선택하는 '오라클'이지만, 이는 실행 간의 노이즈로 인해 과장된 값입니다. 동일한 관찰 방식과 작업을 다시 실행하면 결과의 12-14%가 변경되므로, 이미 보유하고 있는 관찰 방식을 한 번 더 사용하는 것과 새로운 관찰 방식을 추가하는 효과가 거의 같습니다. 중요한 것은 비용 경계입니다. 어떤 관찰 방식으로도 해결할 수 없는 작업만을 가장 저렴한 방식으로 처리하면, 8개의 모든 셀에서 성공률을 유지하면서 비용을 9.5-30.6% 절감할 수 있습니다. 그런 다음 다섯 가지 라우팅 정책(관찰 방식 선택, 강력한 방식을 사용하는 시점 결정, 작업 텍스트를 기반으로 하는 제로 코스트 규칙, 신뢰도 연쇄, 통합된 비용 계층)을 테스트했지만, 어느 것도 단순히 잘 선택된 관찰 방식을 고정하는 것보다 꾸준히 우수한 성능을 보이지 않았습니다. 예외는 가장 희소한 셀에서 나타난 불안정한 결과입니다. 핵심적인 문제는 라우팅에 필요한 감독 정보가 에이전트의 성공률에 의해 결정된다는 점입니다. 즉, 에이전트의 성능이 낮을수록 라우터가 얻는 레이블 수가 줄어들고, 이는 라우팅이 가장 유용할 때 발생하는 현상입니다. 이러한 제한은 현재의 에이전트에 해당되는 문제이며, 라우팅 자체에는 적용되지 않습니다. 레이블 공급과 라우팅 기회는 함께 증가합니다 (셀 간 상관관계 0.95), 따라서 더 강력한 에이전트는 이러한 결과를 뒤집을 수 있으며, 본 연구에서는 재실행 노이즈 범위와 전체 측정 프로토콜을 보고합니다.
Web agents observe a browser through text, pixels, or both, and the choice is usually fixed once for all tasks. We measure six observation modes across eight site-model combinations (cells) on VisualWebArena and WebArena and ask what choosing per task would buy. The modes are complementary: each solves tasks the others miss, they fail in structurally different ways, and the best choice reverses between task sets. The obvious prize, an oracle that picks a winning mode for every task, looks large but is inflated by run-to-run noise: rerunning the same mode on the same tasks changes 12-14% of outcomes, so a second run of a mode already in hand gains about as much as adding a new one. What survives is a cost bound: sending only the tasks no mode solves to the cheapest mode cuts cost by 9.5-30.6% in 8 of 8 cells at unchanged success. We then test five routing policies (picking the mode, deciding when to spend on the strong mode, a zero-cost rule read off the task text, a confidence cascade, and pooled cost tiers), and none robustly beats simply fixing one well-chosen mode; the one exception is a fragile result in our sparsest cell. The central obstruction is that routing supervision is produced at the agent's success rate: the weaker the agent, the fewer labels a router gets, exactly where routing would be most valuable. This limit belongs to today's agents rather than to routing itself. Label supply and routing opportunity rise together (correlation 0.95 across cells), so a stronger agent can overturn the result, and we report the rerun noise bands and the full measurement protocol.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.