MERIT 피드백이 LLM 협상 시스템의 협상 능력을 향상시킨다
MERIT Feedback Elicits Better Bargaining in LLM Negotiators
협상은 종종 논리적인 영역으로 간주되지만, 대규모 언어 모델(LLM)은 제한적인 전략적 깊이와 복잡한 인간 요인에 대한 적응 어려움으로 인해 협상을 효과적으로 수행하는 데 어려움을 겪습니다. 현재의 벤치마크는 이러한 한계를 제대로 반영하지 못합니다. 이러한 격차를 해소하기 위해, 우리는 효용 기반 피드백 중심의 프레임워크를 제시합니다. 우리의 주요 기여는 다음과 같습니다. (i) 다양한 전략 모델링을 지원하는 아홉 가지의 어려운 시나리오(예: 기만, 독점)를 포함하는 새로운 벤치마크인 AgoraBench; (ii) 효용 이론에서 파생된 인간 중심적이고 경제학적으로 타당한 지표. 이는 에이전트 효용, 협상력, 획득 비율을 통해 협상이 인간의 선호도와 얼마나 일치하는지를 간접적으로 측정합니다. (iii) 인간의 선호도를 기반으로 한 데이터셋과 학습 파이프라인을 통해 LLM의 협상 능력을 프롬프팅과 미세 조정 모두를 통해 강화합니다. 실험 결과는 기본 LLM 전략이 종종 인간의 선호도와 다르고, 우리의 메커니즘이 협상 성능을 크게 향상시켜 더 깊은 전략적 행동과 강력한 상대방 인식을 가능하게 한다는 것을 보여줍니다.
Bargaining is often regarded as a logical arena rather than an art or a matter of intuition, yet Large Language Models (LLMs) still struggle to navigate it due to limited strategic depth and difficulty adapting to complex human factors. Current benchmarks rarely capture this limitation. To bridge this gap, we present an utility feedback centric framework. Our contributions are: (i) AgoraBench, a new benchmark spanning nine challenging settings (e.g., deception, monopoly) that supports diverse strategy modeling; (ii) human-aligned, economically grounded metrics derived from utility theory. This is operationalized via agent utility, negotiation power, and acquisition ratio that implicitly measure how well the negotiation aligns with human preference and (iii) a human preference grounded dataset with learning pipeline that strengthens LLMs' bargaining ability through both prompting and finetuning. Empirical results indicate that baseline LLM strategies often diverge from human preferences, while our mechanism substantially improves negotiation performance, yielding deeper strategic behavior and stronger opponent awareness.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.