2607.23991v1 Jul 27, 2026 cs.CL

SyRuP: 보상 기반 예측을 통한 LLM 디코딩 과정에서 시스템 프롬프트 준수성 향상

SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding

Jaehyung Kim
Jaehyung Kim
Citations: 0
h-index: 0
Seoyeon Kim
Seoyeon Kim
Citations: 33
h-index: 2
Minjae Kang
Minjae Kang
Citations: 18
h-index: 2

대규모 언어 모델(LLM)은 역할, 스타일, 형식 및 안전 요구 사항을 지정하는 시스템 프롬프트를 통해 점점 더 많이 제어됩니다. 그러나 모델은 문맥 학습을 통해서만 이러한 프롬프트를 암묵적으로 따르는데, 이는 복잡하거나 조합된 프롬프트의 경우 충분하지 않을 수 있습니다. 기존 접근 방식은 종종 모델 튜닝 또는 응답 수준 재정렬을 필요로 하여 경량 추론 시간 제어에 대한 실용성을 제한합니다. 본 논문에서는 기본 LM을 고정한 상태에서 시스템 프롬프트 준수성을 향상시키는 디코딩 시간 프레임워크인 SyRuP를 소개합니다. SyRuP는 시스템 프롬프트 기반 선호 쌍으로부터 교차 어텐션 보상 헤드를 학습하여, 시스템 프롬프트를 별도의 메모리로 간주하고 토큰 수준의 준수성 점수를 생성합니다. 추론 시, SyRuP는 기본 LM에서 생성된 상위 k개 후보를 재정렬하며, 이때 기본 로짓에 학습된 보상 신호와 시스템으로 인해 발생하는 로짓 변화를 포착하는 선택적 대비 신호를 결합합니다. 시스템 프롬프트 준수성 벤치마크에서의 실험 결과, SyRuP는 일관되게 프롬프팅 및 디코딩 시간 기반 모델보다 우수한 성능을 보이며, 적당한 추론 오버헤드를 갖습니다. 이러한 결과는 명시적인 토큰 수준의 지도가 안정적인 시스템 프롬프트 준수를 위한 효과적이고 실용적인 메커니즘임을 시사합니다.

Original Abstract

Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, styles, formats, and safety requirements. However, models follow these prompts only implicitly through in-context learning, which can be insufficient for complex or compositional prompts. Existing approaches often require model tuning or response-level reranking, limiting their practicality for lightweight inference-time control. We introduce SyRuP, a decoding-time framework for improving system-prompt adherence while keeping the base LM frozen. SyRuP trains a cross-attention reward head from system-prompt-conditioned preference pairs, treating the system prompt as a separate memory to produce token-level adherence scores. At inference, SyRuP reranks the base LM's top-k candidates by combining base logits with the learned reward signal and an optional contrastive signal capturing system-induced logit shifts. Experiments on system-prompt following benchmarks show that SyRuP consistently outperforms prompting and decoding-time baselines with moderate inference overhead. These results suggest that explicit token-level guidance is an effective and practical mechanism for reliable system-prompt following.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!