EgoAfford: 자아 참조 분할을 통한 작업 지향적 기능적 특징 연결
EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation
부분 수준의 기능적 특징 연결은 기본 동작과 관련된 기능적인 객체 영역을 특정하는 데 기여해 왔습니다. 이 능력을 복잡한 작업에 확장하기 위해서는, 관련 객체의 의미론적 역할을 작업 상태와 연관된 시각 정보 및 다단계 계획과 연결해야 합니다. 본 논문에서는 이러한 세 가지 측면을 연결하도록 설계된 벤치마크인 EgoAfford를 소개합니다. EgoAfford는 자아 참조 관찰과 고수준 테이블탑 작업을 입력으로 받아, 모델이 나머지 계획을 생성하고 다음 동작의 최대 세 구성 요소(직접 대상, 도구, 목적지)의 기능적인 영역을 분할하도록 합니다. EgoAfford는 2,000개의 생성된 다단계 장면에서 추출한 약 15,500장의 인간 검증 이미지를 포함하며, 의미론적으로 정렬되고 작업이 완료된 이미지 시퀀스로 구성되어 있습니다. 또한, 26개의 작업에 걸쳐 수동으로 촬영한 102장의 이미지인 EgoAfford-Real 데이터셋도 제공됩니다. 본 논문에서는 역할별 마스크 디코더를 갖춘 3B 멀티모달 대규모 언어 모델인 EgoLens를 이 공동 작업을 위한 도메인 내 참조 모델로 제시합니다. 최근의 참조 분할 LLM, 상용 VLM-SAM2 파이프라인 및 EgoLens에 대한 평가는 다음 단계 추론과 동작 역할 기반 부분 연결의 상호 보완적인 과제를 강조합니다. EgoLens는 생성된 이미지와 수동으로 촬영된 이미지 모두에서 강력한 참조 성능을 보여줍니다. 전체적으로, EgoAfford와 EgoLens는 다단계 테이블탑 작업에서의 인지 및 계획 연구를 위한 기반을 제공합니다. 프로젝트 페이지는 다음 주소에서 확인할 수 있습니다: https://egoafford.github.io
Part-level affordance grounding has advanced the localization of functional object regions associated with elemental actions. Extending this capability to complex tasks calls for connecting the semantic roles of participating objects with task-state-aligned visual observations and multi-step planning. We introduce EgoAfford, a benchmark designed to connect these three aspects. Given an egocentric observation and a high-level tabletop task, a model must generate the remaining plan and segment the functional regions of up to three components of the next action: the direct object, instrument, and destination. EgoAfford comprises approximately 15.5k human-verified images from 2,000 generated multi-step scenes, organized as semantically aligned, task-complete image series, together with EgoAfford-Real, 102 manually captured images spanning 26 tasks. We further present EgoLens, a 3B multimodal large language model with role-specific mask decoders, as an in-domain reference model for this joint task. Evaluations of recent referring-segmentation MLLMs, commercial-VLM--SAM2 pipelines, and EgoLens highlight the complementary challenges of next-step inference and action-role-conditioned part grounding. EgoLens establishes strong reference performance on both generated and manually captured observations. Together, EgoAfford and EgoLens provide a foundation for jointly studying perception and planning in multi-step tabletop tasks. Our project page is available at: https://egoafford.github.io
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.