크로네커 분해 근사 곡률을 통한 태스크 연산에서의 무데이터 가중치 얽힘 해소
Dataless Weight Disentanglement in Task Arithmetic via Kronecker-Factored Approximate Curvature
태스크 연산(Task Arithmetic)은 파운데이션 모델을 적응시키기 위한 모듈식의 확장 가능한 방법을 제공한다. 그러나 여러 태스크 벡터를 결합하면 태스크 간 간섭이 발생하여 표현 표류(representation drift)와 성능 저하를 초래할 수 있다. 표현 표류 정규화는 태스크 벡터의 얽힘을 해소하는 자연스러운 해결책을 제공하지만, 기존 접근법들은 일반적으로 외부 태스크 데이터를 필요로 하여 모듈성 및 데이터 가용성 제약(예: 개인정보 보호 요구사항)과 상충된다. 우리는 표현 표류에 대한 정규화를 곡률 행렬 근사 문제로 공식화하여 데이터가 필요 없는(dataless) 접근법을 제안한다. 이를 통해 잘 확립된 기존 기법들을 활용할 수 있으며, 특히 크로네커 분해 근사 곡률(Kronecker-Factored Approximate Curvature)을 채택하여 태스크 덧셈 및 부정(negation) 작업에서 최고 수준(state-of-the-art)의 성과를 달성하는 실용적인 정규화기를 도출했다. 제안하는 방법은 태스크 수에 대해 일정한 복잡도를 가지며, 태스크 벡터 스케일 재조정에 대한 강건성을 촉진하여 별도의 검증 데이터 튜닝(held-out tuning)의 필요성을 없앤다.
Task Arithmetic yields a modular, scalable way to adapt foundation models. Combining multiple task vectors, however, can lead to cross-task interference, causing representation drift and degraded performance. Representation drift regularization provides a natural remedy to disentangle task vectors; however, existing approaches typically require external task data, conflicting with modularity and data availability constraints (e.g., privacy requirements). We propose a dataless approach by framing regularization against representation drift as a curvature matrix approximation problem. This allows us to leverage well-established techniques; in particular, we adopt Kronecker-Factored Approximate Curvature and obtain a practical regularizer that achieves state-of-the-art results in task addition and negation. Our method has constant complexity in the number of tasks and promotes robustness to task vector rescaling, eliminating the need for held-out tuning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.