2608.02171v1 Aug 03, 2026 cs.AI

프로파일링에서 합성으로: 개인화된 LLM 에이전트의 암묵적 행동 정렬 성능 평가

From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents

Min Zhang
Min Zhang
Citations: 24
h-index: 3
Mong Li Lee
Mong Li Lee
Citations: 2,935
h-index: 14
Zibo Ji
Zibo Ji
Citations: 2
h-index: 1
Hao Fei
Hao Fei
Citations: 23
h-index: 3
Wynne Hsu
Wynne Hsu
Citations: 54
h-index: 2
Bobo Li
Bobo Li
Citations: 1,098
h-index: 17
Jiajia Song
Jiajia Song
Citations: 0
h-index: 0
Haiwen Yi
Haiwen Yi
Citations: 0
h-index: 0
Meishan Zhang
Meishan Zhang
Citations: 8
h-index: 2

대규모 언어 모델(LLM)은 점점 더 강력한 자율 에이전트를 가능하게 하지만, 실용적인 유용성을 확보하기 위해서는 개인화가 여전히 중요합니다. 최근에는 에이전트의 개인화 기능을 평가하는 벤치마크들이 등장했지만, 대부분 정적 선호도 스냅샷, 고정된 상호 작용 로그 또는 미리 정의된 사용자 프로필에 대한 질의응답을 기반으로 합니다. 이러한 설계는 진화하는 사용자의 선호도의 복잡성을 포착하지 못하며, 선호도에 따라 달라지는 작업 실행을 간과한다는 점에서 '지식-행동 격차'라는 문제를 야기합니다. 이 문제를 해결하기 위해, 우리는 노이즈, 암묵적 신호 및 시간적 불일치를 포함하는 장기간의 상호 작용 기록에서 구축된 암묵적 행동 정렬 벤치마크인 IBA-Bench를 소개합니다. 기존 연구와 달리, IBA-Bench는 에이전트가 과거 상호 작용으로부터 추론된 암묵적인 사용자 제약을 만족시키면서 작업을 수행할 수 있는지를 평가합니다. 또한, 우리는 광범위한 검색과 트래jectory 수준의 정렬을 통해 충돌하는 우선순위를 조정하는 에이전트 프레임워크인 IBA-Agent를 제안합니다. IBA-Bench에서의 실험 결과는 최첨단 LLM 에이전트에 대한 효과적인 개인화가 여전히 중요한 과제임을 보여주며, 제안된 IBA-Agent는 9가지 애플리케이션 도메인의 복잡한 시나리오에서 행동 정렬을 크게 향상시킵니다.

Original Abstract

Large Language Models have enabled increasingly capable autonomous agents, yet personalization remains critical for making such agents practically useful. Recent benchmarks have begun evaluating personalization in agents, but they largely rely on static preference snapshots, fixed interaction logs, or question answering over predefined user profiles. Such designs fail to capture the complexity of evolving user preferences and neglect preference-conditioned task execution-a discrepancy we term as the knowledge-to-action gap. To address this challenge, we introduce IBA-Bench, a benchmark for implicit behavioral alignment constructed from longitudinal interaction histories that contain noise, implicit cues, and temporal inconsistencies. Unlike prior work, IBA-Bench evaluates whether an agent can execute tasks while satisfying implicit user constraints inferred from historical interactions. We further propose IBA-Agent, an agent framework that reconciles conflicting priorities through broad retrieval and trajectory-level alignment. Experiment results on IBA-Bench show that effective personalization remains a significant challenge for state-of-the-art LLM agents, and the proposed IBA-Agent substantially improves behavioral alignment in complex scenarios across nine application domains.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!