2602.04863v1 Feb 04, 2026 cs.LG

데이터 속의 잠재적 영향: 로그 선형성을 통한 일반적인 메커니즘

Subliminal Effects in Your Data: A General Mechanism via Log-Linearity

Ishaq Aden-Ali
Ishaq Aden-Ali
Citations: 147
h-index: 6
Noah Golowich
Noah Golowich
Citations: 2,767
h-index: 21
A. Liu
A. Liu
Citations: 17
h-index: 3
Abhishek Shetty
Abhishek Shetty
Citations: 56
h-index: 5
A. Moitra
A. Moitra
Citations: 73
h-index: 4
Nika Haghtalab
Nika Haghtalab
Citations: 4,034
h-index: 26

최근 대규모 언어 모델(LLM) 훈련은 특정 행동을 유도하기 위해 설계된 다양한 알고리즘과 데이터 세트의 집합체가 되었으며, 따라서 데이터 세트가 모델의 특성에 미치는 영향을 이해하는 기술을 개발하는 것이 매우 중요합니다. 최근 실험 결과, 데이터 세트가 개별 데이터 포인트로부터 직접적으로 관찰할 수 없는 신호를 전달할 수 있다는 사실이 밝혀지면서, LLM 훈련에 대한 데이터 중심적인 이해 방식에 근본적인 과제를 제시하고 있습니다. 이러한 현상을 이해하기 위해, 최근 LLM의 선형 구조에 대한 연구에서 영감을 받아, 일반적인 데이터 세트에서 잠재적인 숨겨진 의미가 발생하는 일반적인 메커니즘을 밝혀냈습니다. 저희는 Logit-Linear-Selection (LLS)이라는 방법을 제안합니다. 이는 일반적인 선호 데이터 세트의 부분 집합을 선택하여 다양한 잠재적인 효과를 유도하는 방법을 제시합니다. LLS를 사용하여 실제 데이터 세트의 부분 집합을 발견함으로써, 해당 부분 집합으로 훈련된 모델이 특정 선호도를 갖거나, 데이터 세트에 없는 다른 언어로 응답하거나, 다른 페르소나를 갖는 등 다양한 행동을 보이도록 합니다. 중요한 점은 이러한 효과가 선택된 부분 집합에 대해 유지되며, 다양한 아키텍처를 가진 모델에서도 나타나므로, 그 일반성과 보편성을 뒷받침합니다.

Original Abstract

Training modern large language models (LLMs) has become a veritable smorgasbord of algorithms and datasets designed to elicit particular behaviors, making it critical to develop techniques to understand the effects of datasets on the model's properties. This is exacerbated by recent experiments that show datasets can transmit signals that are not directly observable from individual datapoints, posing a conceptual challenge for dataset-centric understandings of LLM training and suggesting a missing fundamental account of such phenomena. Towards understanding such effects, inspired by recent work on the linear structure of LLMs, we uncover a general mechanism through which hidden subtexts can arise in generic datasets. We introduce Logit-Linear-Selection (LLS), a method that prescribes how to select subsets of a generic preference dataset to elicit a wide range of hidden effects. We apply LLS to discover subsets of real-world datasets so that models trained on them exhibit behaviors ranging from having specific preferences, to responding to prompts in a different language not present in the dataset, to taking on a different persona. Crucially, the effect persists for the selected subset, across models with varying architectures, supporting its generality and universality.

5 Citations
2 Influential
13 Altmetric
74.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!