UESF-Bench: 통합된 목표 추적 및 따라가기 성능 평가를 위한 벤치마크
UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following
언어 지침에 따른 인간 따라가기는 로봇 에이전트에게 중요한 능력입니다. 그러나 기존의 벤치마크는 일반적으로 대상인이 에피소드의 시작 시점에 항상 보이는 것을 전제로 합니다. 이러한 설정은 문제를 단순화하고, 실제 환경에서 에이전트가 언어로 설명된 대상을 먼저 찾아야 하고, 이후 동적인 환경에서 지속적으로 해당 대상을 따라가는 필수 요구 사항을 간과합니다. 최근 연구에서는 인간의 탐색에 대한 연구가 시작되었지만, 기존 벤치마크는 일반적으로 특정 작업 중심 시나리오에서 평가되며, 종종 환경에 대한 더 강력한 사전 지식을 필요로 합니다. 또한, 대부분의 경우 탐색과 따라가기를 별개의 작업으로 취급하며, 체계적인 평가를 위한 통합된 벤치마크가 여전히 부족합니다. 이러한 한계를 해결하기 위해, 우리는 대규모이고 다양한 데이터셋을 갖춘 통합 목표 추적 및 따라가기 벤치마크인 UESF-Bench (Unified Embodied Seeking and Following Benchmark)를 소개합니다. 이 벤치마크는 에이전트가 의미 기반 탐색, 안정적인 행동 전환 및 복구, 그리고 지연된 대상 인식 기능을 수행하도록 요구합니다. 이를 위해, 우리는 잠재 단계 추론과 탐색에서 따라가기로의 전환 모델링을 위한 작업 중심 라우팅 메커니즘을 갖춘 시각-언어-행동 프레임워크인 SeekFollow-VLA를 제안합니다. 실험 결과는 SeekFollow-VLA가 단일 헤드 및 이중 헤드 기반 모델에 비해 단일 인물 환경과 다중 인물 환경 모두에서 상당한 성능 향상을 보여주며, 통합된 목표 추적 및 따라가기 시스템의 기준을 제시함을 나타냅니다.
Language-guided human following is an important capability for embodied agents, but existing benchmarks typically assume that the target person is visible at the start of an episode. This setting simplifies the problem and overlooks a more realistic requirement: an agent often needs to first find a language-described target and then persistently follow that target in a dynamic environment. While recent work has started to study human search, existing settings are typically evaluated in task-specific scenarios and often rely on stronger prior knowledge of the environment. Moreover, they usually treat searching and following as separate tasks and still lack a unified benchmark for systematic evaluation. To address these limitations, we introduce the Unified Embodied Seeking and Following Benchmark (UESF-Bench), a large-scale and diverse benchmark for embodied human seeking and following. The benchmark requires agents to handle semantic-guided exploration, reliable behavior switching and recovery, and delayed identity grounding. To this end, we propose SeekFollow-VLA, a vision-language-action framework with a task-driven routing mechanism for latent phase inference and transition modeling between seeking and following. Experimental results show that SeekFollow-VLA achieves clear improvements over both single-head and dual-head baselines across single-person and multi-person environments, establishing a baseline for unified embodied seek-and-follow.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.