작업 벡터가 언제 간섭하는가? 가중치 공간 조합의 유효성 경계를 분석하다
When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space Composition
작업 연산(task arithmetic)은 미세 조정으로 인한 변화를 가중치 공간에서 합쳐지는 방향으로 취급하지만, 파라미터 추가가 모델 기능에 예측 가능한 변화를 가져오는지 여부는 아직 명확하지 않습니다. 본 연구에서는 파라미터 기하학적 구조와 기능적 기하학적 구조를 분리하고, 입력 분포에 조건화된 첫 번째 토큰 예측 분포 상호 작용 비율을 사용하여 2차원 작업 벡터 표면에서 두 항목 간의 기능적 비가산성을 측정합니다. 이때 정규화된 제어 그룹, 세 가지 학습 시드, 그리고 응답 기반 미세 조정을 사용했습니다. Qwen2.5-1.5B 모델에서 코드와 안전성 관련 작업은 코드 및 명령어 프롬프트에 대해 일치된 코드+수학 제어 그룹보다 더 큰 비가산성을 보였지만, 수학 프롬프트에서는 그렇지 않았습니다. 사전에 정의된 6개의 작업 조합을 확장했을 때, 관찰되지 않은 작업 쌍의 8가지 고(high) 대비 저(low) 비교에서 모두 예측된 방향이 나타났습니다. 이러한 주요 순서는 0.5B 규모의 전체 파라미터 미세 조정, Qwen2.5 LoRA를 활용한 최대 7B 규모 테스트, 그리고 Llama-3.1-8B 모델 아키텍처 검증을 통해 지속적으로 확인되었습니다. 외부 검증 결과, 원본 공개 코드, 명령어 및 안전성 프롬프트는 이러한 연속적인 대비를 유지하지만, 명령어 스타일 래퍼는 동일한 공개 코드 프롬프트에서 이를 무너뜨리고, EvalPlus pass@1 상호 작용은 이를 안정적으로 재현하지 못합니다. 따라서 가중치 공간 조합은 다양한 적응 방법, 규모 및 추가 모델 패밀리를 통해 입력 및 형식에 조건화된 기능적 관계를 뒷받침하지만, 보편적인 성능 예측 도구는 아닙니다.
Task arithmetic treats fine-tuning displacements as composable directions in weight space, yet it remains unclear when parameter addition reflects predictable changes in model function. We separate parameter geometry from functional geometry and measure pairwise functional non-additivity over a two-dimensional task-vector surface, using a first-token predictive-distribution interaction ratio conditioned on an input distribution and evaluated with norm-matched controls, three training seeds, and response-only fine-tuning. On Qwen2.5-1.5B, code+safety is more non-additive than the matched code+math control on code and instruction prompts, but not on math prompts. In a prospectively specified six-task expansion, all eight high-versus-low comparisons of unseen task pairs have the predicted sign. The primary ordering further persists under full-parameter fine-tuning at 0.5B, Qwen2.5 LoRA scale tests up to 7B, and a Llama-3.1-8B cross-architecture audit. External validation exposes a sharper boundary: raw public code, instruction, and safety prompts preserve the continuous contrast, whereas an instruction-style wrapper collapses it on the identical public-code prompts, and EvalPlus pass@1 interactions do not robustly reproduce it. Weight-space composition therefore supports coarse, input- and format-conditioned functional statements across adaptation methods, scales, and one additional model family, not a universal merging-performance predictor.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.