2608.04339v1 Aug 05, 2026 cs.CL

제약 조건 기반 혼합 전략 그룹 DRO를 이용한 공정한 시스템 프롬프트 선택

Equitable System-Prompt Selection via Constrained Mixed-Strategy GroupDRO

C. Gao
C. Gao
Citations: 10
h-index: 2
Kezhen Chen
Kezhen Chen
Citations: 244
h-index: 4
Ruiyao Xu
Ruiyao Xu
Citations: 104
h-index: 3
Qiaoxin Yang
Qiaoxin Yang
Citations: 1
h-index: 1
Mengyu Xu
Mengyu Xu
Citations: 0
h-index: 0
Zhihan Liu
Zhihan Liu
Citations: 0
h-index: 0
Zachary Liu
Zachary Liu
Citations: 90
h-index: 2

대규모 언어 모델은 정보 검색에 점점 더 많이 사용되고 있지만, 의미적으로 동일하지만 표현 방식이 다른 질문에 대해 현저히 다른 품질의 답변을 얻을 수 있습니다. 시스템 프롬프트는 응답 동작을 제어하는 데 널리 사용되지만, 일반적으로 평균적인 품질을 기준으로 최적화되기 때문에 특정 질문 표현은 여전히 불완전하거나 낮은 품질의 답변을 받을 수 있습니다. 이를 해결하기 위해, 우리는 시스템 프롬프트 선택을 위한 제약 조건 기반 혼합 전략 그룹 DRO(GroupDRO) 프레임워크를 제안합니다. 이 프레임워크는 기존에 존재하는 시스템 프롬프트 풀 내에서 각 시스템 프롬프트에 가중치를 부여하여 평가 지표 및 그룹 전반에 걸쳐 최악의 정보 품질 손실을 최소화하고, 동시에 평균 기반 선택과 유사한 수준으로 평균 손실이 유지되도록 합니다. 풀 생성 및 선택이 분리되어 있어, 이 방법은 모든 시스템 프롬프트 풀에 적용될 수 있으며 단일 프롬프트 대신 상호 보완적인 여러 프롬프트를 활용할 수 있습니다. 두 개의 이중 언어 의료 및 소비자 금융 벤치마크에서 다섯 가지 LLM을 사용하여 실험한 결과, 제안된 방법은 완화되지 않은 경우보다 평균 전체 성능(Overall Mean), 최악의 25% 성능(Worst 25% Mean), 그리고 최악의 성능(Worst)이 각각 13.1%, 13.2%, 13.7% 감소했습니다. 또한, 전반적인 품질은 평균 선택과 유사하게 유지되었습니다. 제안된 방법의 다중 프롬프트 가중치는 지표-그룹 쌍 간의 상호 보완성을 보여줍니다. 코드 및 데이터는 다음 링크에서 확인할 수 있습니다: https://github.com/Rainxu09/equitable-system-prompt-selection.

Original Abstract

Large language models are increasingly used for information seeking, yet semantically equivalent questions phrased in different ways can receive answers of considerably different quality. System prompts are widely employed to steer response behavior, but they are typically optimized for average-case quality, so some question phrasings may still receive incomplete or low-quality answers. To address this, we formulate a constrained mixed-strategy GroupDRO framework for system-prompt selection. Instead of optimizing the system-prompt text, the framework assigns weights to system prompts in an existing pool to minimize the worst-case information-quality loss across evaluation metrics and groups, while constraining the mean loss to stay close to that of average-based selection. Because pool generation and selection are decoupled, the method applies to any system-prompt pool and can leverage an ensemble of complementary system prompts rather than a single one. Across five LLMs on two bilingual medical and consumer-finance benchmarks, the constrained method reduces the Overall Mean, Worst 25% Mean, and Worst by 13.1%, 13.2%, and 13.7% on average relative to no mitigation while keeping overall quality close to Average selection. Its multi-prompt weights reveal complementarity across metric-group pairs. Code and data are available at https://github.com/Rainxu09/equitable-system-prompt-selection.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!