2606.18847v1 Jun 17, 2026 cs.AI

WorldLines: 장기 상태 기반 임베디드 에이전트의 성능 측정 및 모델링

WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents

Zexi Li
Zexi Li
Citations: 103
h-index: 3
Yinchuan Li
Yinchuan Li
Citations: 94
h-index: 5
Yifan Chang
Yifan Chang
Citations: 26
h-index: 2
Xinli Xu
Xinli Xu
Citations: 73
h-index: 6
Yehang Zhang
Yehang Zhang
Citations: 12
h-index: 2
Jia Su
Jia Su
Citations: 11
h-index: 1
Haojian Huang
Haojian Huang
Citations: 3
h-index: 1
T. Zhou
T. Zhou
Citations: 3,287
h-index: 2
Yingjie Xu
Yingjie Xu
Citations: 126
h-index: 4
Ying-Cong Chen
Ying-Cong Chen
Citations: 40
h-index: 3

인간을 오랫동안 지원하기 위해 실제 가정 환경에서 사용되는 임베디드 에이전트는 사용자 루틴, 세계 상태 및 과거 상호 작용을 기억해야 합니다. 기존의 장기 기억 벤치마크는 주로 언어 기반 검색 및 질문 답변을 평가하는 데 중점을 두는 반면, 임베디드 벤치마크는 종종 짧은 시간 범위의 작업 실행에 초점을 맞추고 동적 환경에서 장기 기억 사용 여부를 테스트하지 않습니다. 본 연구에서는 장기적인 가정 지원을 위한 프로젝트 기반 벤치마크인 WorldLines를 소개합니다. WorldLines는 대화, 행동, 실행 피드백, 객체 및 장치 상태 변경 사항으로 구성된 시간적으로 확장된 가정 환경 데이터를 구축하고, 이를 메모리 질문 답변(Memory QA)과 임베디드 작업 계획을 위한 증거 연결 샘플로 변환합니다. 또한, 가시성을 고려한 기억 관리와 액션 기반의 상태 추적을 통해 상태 인지적인 의사 결정을 지원하는 관찰자 기반 기억 프레임워크인 ObsMem을 제안합니다. 실험 결과는 부분 관찰성, 덮어쓰여진 세계 상태 및 장기 기억을 임베디드 계획으로 변환하는 데 있어 여전히 많은 어려움이 있음을 보여주며, ObsMem은 이러한 환경에 대한 더욱 강력한 참조 아키텍처를 제공합니다.

Original Abstract

To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and past interactions. Existing long-term memory benchmarks mainly evaluate language-centric retrieval and question answering, while embodied benchmarks often focus on short-horizon task execution without testing long-term memory use in dynamic environments. We introduce WorldLines, a project-driven benchmark for long-horizon embodied household assistance. It constructs temporally extended household traces with dialogues, actions, execution feedback, object and device state changes, and converts them into evidence-linked samples for Memory QA and Embodied Task Planning. We further propose ObsMem, an observer-grounded memory framework that maintains visibility-aware memories and action-native state trails for state-aware decisions. Experiments reveal persistent challenges in partial observability, overwritten world states, and translating long-term memory into embodied plans, while ObsMem offers a stronger reference architecture for this setting.

1 Citations
0 Influential
3 Altmetric
16.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!