2603.04045v1 Mar 04, 2026 cs.LG

단백질 언어 모델에서의 추론 시 독성 완화: 도메인 적응과 추론 시간 제어

Inference-Time Toxicity Mitigation in Protein Language Models

Manuel Fern'andez Burda
Manuel Fern'andez Burda
Citations: 6
h-index: 1
Santiago Aranguri
Santiago Aranguri
Citations: 23
h-index: 2
Iván Arcuschin Moreno
Iván Arcuschin Moreno
Citations: 6
h-index: 2
Enzo Ferrante
Enzo Ferrante
Citations: 8
h-index: 2

단백질 언어 모델(PLM)은 새로운 단백질 설계에 유용한 도구으로 자리 잡고 있지만, 이중 용도 가능성은 안전 문제를 야기합니다. 본 연구에서는 특정 분류군에 대한 도메인 적응이 독성 단백질 생성으로 이어질 수 있으며, 이는 독성이 학습 목표가 아님에도 발생할 수 있음을 보여줍니다. 이러한 문제를 해결하기 위해, 본 연구는 PLM의 추론 시간 제어 메커니즘으로 로짓 차이 증폭(Logit Diff Amplification, LDA)을 적용했습니다. LDA는 재학습 없이, 기준 모델과 독성-미세 조정 모델 간의 로짓 차이를 증폭하여 토큰 확률을 수정합니다. 네 가지 분류군에 대해 LDA는 일관되게 ToxDL2를 통해 측정된 예측 독성률을 독성-미세 조정 기준선보다 낮추면서도 생물학적 타당성을 유지합니다. Fréchet ESM 거리와 예측 폴딩 가능성(pLDDT)을 사용하여 품질을 평가한 결과, LDA는 천연 단백질과의 분포적 유사성을 유지하고 구조적 생존성을 유지하는 것으로 나타났습니다(활성화 기반 제어 방법은 서열 특성을 저하시키는 경향이 있습니다). 본 연구 결과는 LDA가 생성 품질을 유지하면서 유도된 독성을 완화하는 단백질 생성기의 실용적인 안전 제어 장치임을 보여줍니다.

Original Abstract

Protein language models (PLMs) are becoming practical tools for de novo protein design, yet their dual-use potential raises safety concerns. We show that domain adaptation to specific taxonomic groups can elicit toxic protein generation, even when toxicity is not the training objective. To address this, we adapt Logit Diff Amplification (LDA) as an inference-time control mechanism for PLMs. LDA modifies token probabilities by amplifying the logit difference between a baseline model and a toxicity-finetuned model, requiring no retraining. Across four taxonomic groups, LDA consistently reduces predicted toxicity rate (measured via ToxDL2) below the taxon-finetuned baseline while preserving biological plausibility. We evaluate quality using Fréchet ESM Distance and predicted foldability (pLDDT), finding that LDA maintains distributional similarity to natural proteins and structural viability (unlike activation-based steering methods that tend to degrade sequence properties). Our results demonstrate that LDA provides a practical safety knob for protein generators that mitigates elicited toxicity while retaining generative quality.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!