2601.05215v2 Jan 08, 2026 cs.AI

MineNPC-Task: 메모리 인식 마인크래프트 에이전트를 위한 태스크 스위트

MineNPC-Task: Task Suite for Memory-Aware Minecraft Agents

Tamil Sudaravan Mohan Doss
Tamil Sudaravan Mohan Doss
Citations: 1
h-index: 1
Andrew D. Wilson
Andrew D. Wilson
Citations: 7
h-index: 1
B. T. Kumaravel
B. T. Kumaravel
Citations: 439
h-index: 9
Michael Xu
Michael Xu
Citations: 106
h-index: 5
Sudha Rao
Sudha Rao
Citations: 148
h-index: 7

우리는 오픈 월드 마인크래프트에서 메모리 인식 및 혼합 주도(mixed-initiative) LLM 에이전트를 테스트하기 위한 사용자 작성 벤치마크이자 평가 도구인 MineNPC-Task를 제안한다. 합성 프롬프트에 의존하는 대신, 작업들은 숙련된 플레이어들과의 형성적 및 총괄적 협동 플레이를 통해 도출되었으며, 이후 명시적인 사전 조건과 종속성 구조를 갖춘 매개변수 템플릿으로 정규화되었다. 이러한 작업들은 게임 세계 밖의 정보를 이용한 편법을 금지하는 제한된 지식 정책 하에서, 기계적으로 확인 가능한 검증기와 짝을 이룬다. 이 평가 도구는 계획 미리보기, 대상별 명확화, 메모리 읽기 및 쓰기, 사전 조건 확인, 수정 시도를 포함한 계획, 행동, 메모리 이벤트를 기록하며, 오직 게임 내 증거만을 사용하여 시도된 총 하위 작업 수 대비 결과를 보고한다. 초기 연구 사례로, 우리는 GPT-4o를 사용하여 프레임워크를 구현하고 8명의 숙련된 플레이어를 대상으로 216개의 하위 작업을 평가했다. 우리는 코드 실행, 인벤토리 및 도구 조작, 참조, 길 찾기에서 반복되는 오류 패턴을 관찰했으며, 동시에 혼합 주도적 설명과 경량 메모리 사용을 통해 성공적으로 복구되는 사례도 확인했다. 참가자들은 상호작용 품질과 인터페이스 사용성을 긍정적으로 평가했으나, 작업 간 더 강력한 메모리 지속성의 필요성을 언급했다. 우리는 향후 메모리 인식 임바디드 에이전트의 투명하고 재현 가능한 평가를 지원하기 위해 전체 태스크 스위트, 검증기, 로그 및 평가 도구를 공개한다.

Original Abstract

We present MineNPC-Task, a user-authored benchmark and evaluation harness for testing memory-aware, mixed-initiative LLM agents in open-world Minecraft. Rather than relying on synthetic prompts, tasks are elicited through formative and summative co-play with expert players, then normalized into parametric templates with explicit preconditions and dependency structure. These tasks are paired with machine-checkable validators under a bounded-knowledge policy that forbids out-of-world shortcuts. The harness captures plan, action, and memory events, including plan previews, targeted clarifications, memory reads and writes, precondition checks, and repair attempts, and reports outcomes relative to the total number of attempted subtasks using only in-world evidence. As an initial snapshot, we instantiate the framework with GPT-4o and evaluate 216 subtasks across 8 experienced players. We observe recurring breakdown patterns in code execution, inventory and tool handling, referencing, and navigation, alongside successful recoveries supported by mixed-initiative clarifications and lightweight memory use. Participants rated interaction quality and interface usability positively, while noting the need for stronger memory persistence across tasks. We release the complete task suite, validators, logs, and evaluation harness to support transparent and reproducible evaluation of future memory-aware embodied agents.

1 Citations
0 Influential
4.5 Altmetric
23.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!