GESTO: 동적 환경에서의 추론을 위한 인간 중심의 시공간 기억 시스템
GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes
인간 환경에서 작동하는 로봇은 단순히 객체가 무엇이고 어디에 있는지 뿐만 아니라, 사람들이 시간이 지남에 따라 이를 어떻게 사용하는지, 그리고 개별적인 상호작용이 어떻게 목표 지향적인 활동으로 구성되는지를 이해할 수 있는 기억 기능이 필요합니다. 기존의 4차원 장면 그래프는 객체와 장소의 역사를 보존하지만 활동 구조를 포함하지 않으며, 반면 활동 표현은 종종 지속적인 3D 장면과 연결되지 않거나 외부에서 제공된 이벤트 경계 및 객체 연관성에 의존합니다. 본 논문에서는 GESTO (Grounded Event and Spatio-Temporal memOry)라는 시공간 기억 시스템을 제안합니다. GESTO는 지속적인 4차원 장면 그래프와 원자적 인간-객체 상호작용과 목표 지향적인 이벤트의 두 수준으로 구성된 계층 구조를 결합합니다. RGB-D 관찰 데이터를 기반으로 GESTO는 타임스탬프가 포함된 상호작용을 자동으로 추출하고, 이를 지속적인 장면 요소에 연결하며, 상호작용들을 이벤트로 그룹화하고, 이벤트 문맥을 사용하여 불확실한 객체 연관성을 개선합니다. 관계 인식 도구 사용 에이전트는 결과적으로 생성된 기억 시스템을 활용하여 활동 중심의 시공간 추론을 수행합니다. GESTO는 기존 벤치마크의 재현 가능한 텍스트, 이진 및 시간 범주에 대한 평가와 함께 40개의 새로운 Space2Event 및 Event2Space 질의를 사용하여 평가되었습니다. GESTO는 표준 범주에서 각각 0.71, 0.75, 0.70의 점수를 달성했으며, 이는 실제 이벤트 및 객체 정보를 제공받은 기존 방법과 유사한 성능을 보입니다. 또한 동일한 추론 프레임워크를 사용할 때 입력 정보가 제거된 경우 상당한 성능 향상을 보여줍니다. 추가적으로 Space2Event 및 Event2Space 질의에서 각각 0.73 및 0.75의 점수를 달성했습니다. 분석 결과, 계층적인 이벤트 구조와 문맥 인식 기반의 연관성 개선은 상호 보완적인 이점을 제공하며, 이는 동적 인간 환경에서의 활동 기반 계층적 기억 시스템을 활용한 후향적 추론에 기여합니다.
Robots operating in human environments need memories that capture not only what objects exist and where, but also how people use them over time and how individual interactions compose into goal-directed activities. Existing 4D scene graphs preserve object and place histories but omit activity structure, whereas activity representations are either not grounded in persistent 3D scenes or rely on externally provided event boundaries and object associations. We present GESTO (Grounded Event and Spatio-Temporal memOry), a spatio-temporal memory that couples a persistent 4D scene graph with a two-level hierarchy of atomic human--object interactions and goal-driven events. From an RGB-D observation stream, GESTO automatically extracts timestamped interactions, grounds them to persistent scene entities, groups them into events, and uses event context to refine uncertain object associations. A relation-aware tool-calling agent queries the resulting memory for activity-centric spatio-temporal reasoning. We evaluate GESTO on the reproducible text, binary, and time categories of an existing benchmark, together with 40 new Space2Event and Event2Space queries. GESTO achieves scores of 0.71, 0.75, and 0.70 on the standard categories, approaching a method supplied with ground-truth event and object grounding, while substantially outperforming the same reasoning framework when these inputs are removed. It further achieves 0.73 and 0.75 on Space2Event and Event2Space queries. Ablations show that hierarchical event structure and context-aware grounding refinement provide complementary benefits, supporting activity-grounded hierarchical memory for retrospective reasoning in dynamic human environments.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.