2605.29816v1 May 28, 2026 cs.AI

대규모 언어 모델에서 적대적이지 않은 강건성 활용

Harnessing non-adversarial robustness in large language models

I. Oseledets
I. Oseledets
Citations: 710
h-index: 13
Mikhail Seleznyov
Mikhail Seleznyov
Citations: 62
h-index: 4
Elena Tutubalina
Elena Tutubalina
Citations: 20
h-index: 2
Alexander Panchenko
Alexander Panchenko
Citations: 503
h-index: 7
Qinghua Zhou
Qinghua Zhou
Citations: 94
h-index: 6
E. Aleshina
E. Aleshina
Citations: 0
h-index: 0
Andrey Lovyagin
Andrey Lovyagin
Citations: 0
h-index: 0
O. Somov
O. Somov
Citations: 0
h-index: 0
Ivan Tyukin
Ivan Tyukin
Citations: 12
h-index: 2

본 연구는 대규모 언어 모델(LLM)의 강건성을 향상시키는 방법을 제시합니다. 특히, 의미적으로 유사하지만 텍스트적으로 다른 프롬프트로 인해 발생하는 변화 및 잠재적인 오류에 대한 LLM의 대응 능력을 개선하는 데 중점을 둡니다. 최근 연구 결과에 따르면, 이러한 종류의 프롬프트 변형은 LLM의 작업 성능에 상당한 영향을 미칠 수 있습니다. 본 연구의 핵심 질문은 다음과 같습니다: 전체 모델을 재학습하지 않고도, 의미적으로 중립적인 프롬프트 변경에 대한 LLM의 강건성을 확보할 수 있는가? 우리는 이론적 분석과 실험을 통해 이 질문에 답하고자 합니다. 우리의 이론적 분석 결과, 모델의 강건성에 영향을 미치는 중요한 요소는 신경망 모듈 출력에서 발생하는 체계적인 기대 편향 또는 교란 유발 편향이라는 것을 밝혀냈습니다. 이러한 분석 결과를 바탕으로, 우리는 간단한 파인튜닝 과정을 통해 강건성을 달성할 수 있음을 보여줍니다: 강건성을 위한 편향 제거(debiasing for robustness). 또한, 어떤 조건에서 편향 제거가 도움이 되고, 어떤 경우에는 그렇지 않은지 확인하고, 이론적 분석과 광범위한 실험을 통해 편향 제거를 통한 강건성 확보가 LLM의 강건성을 향상시키고 무작위 프롬프트 변경에 대한 인증을 제공하는 빠르고 효율적인 방법이 될 수 있음을 입증합니다.

Original Abstract

The work presents an approach for addressing the challenge of robustness in Large Language Models (LLMs) to alterations and potential errors caused by semantically similar but textually different prompts. Recent works have shown that these kinds of prompt variations can significantly impact the performance of LLMs on tasks. The central question is: can LLMs' robustness to semantically-neutral prompt alterations be acquired without expensive retraining of the entire model? We address this question both theoretically and through experiments. Our theoretical analysis reveals a crucial factor impacting model robustness - a systematic expected shift or perturbation-induced bias in neural network module outputs. Motivated by this analysis, we show that robustness can be achieved via a simple fine-tuning process: debiasing for robustness. We identify conditions when debiasing helps and when it does not, and demonstrate, through both theory and extensive experiments, that debiasing for robustness may indeed be a quick and efficient tool to enhance robustness and provide certification against random prompt perturbations.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!