PersonalHomeBench: 개인 맞춤형 스마트 홈 환경에서 에이전트 성능 평가
PersonalHomeBench: Evaluating Agents in Personalized Smart Homes
에이전트 기반 인공지능 시스템은 실제 응용 분야로 빠르게 발전하고 있지만, 복잡하고 개인화된 환경에서의 성능은 아직 충분히 검증되지 않았습니다. 이러한 격차를 해소하기 위해, 개인 맞춤형 스마트 홈 환경에서 기초 모델을 에이전트 지원 시스템으로 평가하기 위한 벤치마크인 PersonalHomeBench를 소개합니다. 이 벤치마크는 반복적인 과정을 통해 풍부한 가정 환경 상태를 구축하고, 이를 바탕으로 개인화되고 상황에 따른 다양한 작업을 생성합니다. 현실적인 에이전트-환경 상호 작용을 지원하기 위해, 가정 정보 검색, 가전 제품 제어, 상황 인지 기능을 제공하는 종합적인 도구 모음인 PersonalHomeTools를 제공합니다. PersonalHomeBench는 단일 모드 및 다중 모드 관찰 환경에서 에이전트의 반응형 및 선제적 능력을 평가합니다. 철저한 실험 결과, 작업 복잡도가 증가함에 따라 체계적인 성능 저하가 나타났으며, 특히 반사실적 추론 및 부분 관찰 환경에서 에이전트는 심각한 오류를 보였습니다. 이러한 결과는 PersonalHomeBench를 개인 맞춤형 에이전트의 추론 및 계획의 견고성 및 한계를 분석하기 위한 엄격한 평가 플랫폼으로 자리매김하게 합니다.
Agentic AI systems are rapidly advancing toward real-world applications, yet their readiness in complex and personalized environments remains insufficiently characterized. To address this gap, we introduce PersonalHomeBench, a benchmark for evaluating foundation models as agentic assistants in personalized smart home environments. The benchmark is constructed through an iterative process that progressively builds rich household states, which are then used to generate personalized, context-dependent tasks. To support realistic agent-environment interaction, we provide PersonalHomeTools, a comprehensive toolbox enabling household information retrieval, appliance control, and situational understanding. PersonalHomeBench evaluates both reactive and proactive agentic abilities under unimodal and multimodal observations. Thorough experimentation reveals a systematic performance reduction as task complexity increases, with pronounced failures in counterfactual reasoning and under partial observability, where effective tool-based information gathering is required. These results position PersonalHomeBench as a rigorous evaluation platform for analyzing the robustness and limitations of personalized agentic reasoning and planning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.