도움의 저주: DistractionIF를 통한 주의 분산 지시문에 대한 강건성에서 나타나는 역 스케일링 법칙
The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions via DistractionIF
대규모 언어 모델(LLM)은 점점 더 많이 에이전트 및 검색 증강 생성(RAG) 시스템에 배포되고 있으며, 여기서 사용자가 지정한 작업을 외부에서 제공된 참조 텍스트를 기반으로 수행해야 합니다. 실제로 이러한 컨텍스트는 종종 구조화되지 않고 편집자 코멘트나 시스템 로그와 같은 무해하지만 지시어와 유사한 의미 잡음이 포함되어 있는데, 이는 엄격하게 데이터로 취급되어야 합니다. 우리는 참조 텍스트에서 그러한 주의 분산 지시에 대한 강건성을 평가하도록 설계된 벤치마크인 DistractionIF를 소개합니다. 다양한 모델을 대상으로 실험한 결과, 일관된 역 스케일링 현상이 관찰되었습니다. 즉, 더 큰 모델은 종종 강건성이 낮으며, 규모가 증가함에 따라 성능이 최대 30포인트까지 감소하는 경향이 있습니다. 메커니즘적으로, 우리의 퍼플렉시티 분석 결과는 확장(scaling)이 강건한 동작과 주의 분산된 동작 사이의 확률적 경계를 약화시켜 모델이 잡음을 지시어로 과도하게 해석할 가능성을 높인다는 것을 보여줍니다. 이를 해결하기 위해, 강화 학습, 특히 그룹 상대 정책 최적화(GRPO)를 통해 이러한 경계를 복원하여 일반적인 지시 따르기 능력에 손상을 주지 않고 강건성을 최대 15.5%까지 향상시킬 수 있음을 보여주었습니다. 우리의 연구 결과는 참조 기반 작업에서 중요한 지시 따르기 강건성 격차가 존재하며, 강화 학습이 대규모로 엄격한 데이터-지시 분리를 구현하는 데 유망한 방법임을 강조합니다.
Large Language Models (LLMs) are increasingly deployed in agentic and retrieval-augmented generation (RAG) systems, where they must execute user-specified tasks over externally provided reference text. In practice, such context is often unstructured and contaminated with benign but instruction-like semantic noise, such as editorial comments and system traces, which should be treated strictly as data. We introduce DistractionIF, a benchmark designed to evaluate robustness against such distractor instructions in reference text. Across a broad range of models, we observe a consistent inverse scaling phenomenon: larger models are often less robust, with performance dropping by up to 30 points as scale increases. Mechanistically, our perplexity analysis reveals that scaling erodes the probabilistic boundary between robust and distracted behaviors, making models increasingly prone to over-interpreting noise as instructions. To address this, we demonstrate that reinforcement learning, specifically Group Relative Policy Optimization (GRPO), can restore this boundary, improving robustness by up to 15.5% without compromising general instruction-following capability. Our findings highlight a critical instruction-following robustness gap in reference-grounded tasks and establish reinforcement learning as a promising path for enforcing strict data-instruction separation at scale.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.