의사 결정 수준 조작: 비트 플립 공격을 통한 대규모 언어 모델에 인지적 편향 주입
Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks
대규모 언어 모델(LLM)은 기업 전략과 같은 중요한 의사 결정 시나리오에서 널리 활용되고 있으며, 사용자들은 점점 더 LLM의 결과물을 의존하고 있습니다. 그러나 오픈 소스 모델 공유 생태계와 LLM 기반의 중요한 의사 결정 애플리케이션 간의 깊은 통합은 심각한 위험을 초래합니다. 공격자가 모델의 인지적 입장을 조작할 수 있다면, 이는 간접적으로 하위 의사 결정자의 판단과 행동에 영향을 미칠 수 있습니다. 본 논문에서는 이러한 위협을 '의사 결정 수준 조작'이라고 정의합니다. 기존 공격은 금지된 콘텐츠를 유발하거나 모델 기능을 저하시키지 않고도 특정 인지적 조작을 달성하는 데 실패합니다. 이 문제를 해결하기 위해, 본 논문에서는 비트 플립 공격(BFA)이 의사 결정 수준 조작을 위한 공격 벡터로 사용될 수 있음을 밝힙니다. BFA는 실시간 상호 작용이나 학습 과정에 대한 제어가 필요 없으며, 배포 후에도 극소수의 가중치 비트만 변경하면 은밀하고 저렴하며 지속적인 인지적 조작이 가능합니다. 따라서, 본 논문에서는 LLM을 위한 인지적 편향 주입 프레임워크인 CogBias를 제안합니다. CogBias는 미분 가능한 감정 평가기를 통해 주관적인 선호도를 최적화 신호로 변환하고, 다중 목적 손실 함수를 사용하여 여러 측면을 동시에 제약하며, BitScout을 구축하여 중요한 비트를 찾아 초소량의 비트 변경으로 표적 인지적 개입을 달성합니다. Llama-3.2-3B, Mistral-7B, Qwen2.5-14B 모델과 상업용 추천 및 논쟁적인 사실 주제 시나리오에 대한 실험 결과, 소수의 비트를 플립하는 것만으로도 표적 주제에 대해 상당한 입장 변화를 안정적으로 유발할 수 있으며, 동시에 대상이 아닌 작업 및 전체 출력 분포에 미치는 영향은 제한적이었습니다. 본 연구는 LLM의 고수준 가치 정렬을 저해하기 위해 소량의 저수준 가중치 데이터 변경만으로도 충분하다는 것을 보여줍니다.
Large Language Models (LLMs) have been widely applied in high-stakes decision-making scenarios such as corporate strategy, and users are increasingly relying on their outputs. However, the deep integration of open-source model sharing ecosystems with LLM-powered critical decision-making applications also introduces critical risks: if an attacker can manipulate the model's cognitive stance, they can indirectly influence the judgments and actions of downstream decision-makers. This paper defines such threats as decision-level hijacking. Existing attacks fail to achieve targeted cognitive manipulation without triggering prohibited content or degrading model functionality. To fill this gap, this paper reveals that Bit-Flip Attacks (BFAs) can serve as an attack vector for inducing decision-level hijacking, requiring no real-time interaction or control over the training process, and only a minimal number of weight bits need to be flipped after deployment to achieve stealthy, low-cost, and persistent cognitive manipulation. Therefore, we propose CogBias, a cognitive bias injection framework for LLMs. CogBias converts subjective preferences into optimization signals via a differentiable sentiment evaluator, uses a multi-objective loss to jointly constrain multiple dimensions, and constructs BitScout to locate critical bits, achieving targeted cognitive intervention under an ultra-sparse flip budget. Experiments on Llama-3.2-3B, Mistral-7B, and Qwen2.5-14B, as well as on the commercial recommendation and controversial factual topic scenarios, demonstrate that flipping only a small number of bits stably induces significant stance shifts on target topics, while the impact on non-target tasks and overall output distribution is limited. This work demonstrates that minute perturbations to low-level weight data suffice to undermine the high-level value alignment of LLMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.