2608.04425v1 Aug 05, 2026 cs.RO

SSC: 양손 조작 작업 라벨링을 위한 검증 가능한 구조화된 표현

SSC: A Verifiable Structured Representation for Bimanual Manipulation Labelling

Yupu Lu
Yupu Lu
Citations: 119
h-index: 6
Jia Pan
Jia Pan
Citations: 40
h-index: 3
Shuang Wu
Shuang Wu
Citations: 80
h-index: 2
Ruihua Han
Ruihua Han
Citations: 22
h-index: 2
Marcus Kalander
Marcus Kalander
Citations: 398
h-index: 7
Sihan Chen
Sihan Chen
Citations: 0
h-index: 0
Yi-Chi Zhang
Yi-Chi Zhang
Citations: 0
h-index: 0

부분 작업 레이블은 긴 시간 동안의 조작 시연을 정책 학습 및 평가를 위한 더 짧은 의미 단위로 분해합니다. 자연어 설명은 읽기 쉽지만, 언어적 다양성으로 인해 자동 검증이 어렵습니다. BEHAVIOR-1K의 skill_annotation과 같은 엄격한 템플릿 형식은 언어적으로 과도하게 세분화되어 가독성과 주석 일관성을 저해합니다. 우리는 이러한 극단적인 경계를 해소하는 구조화된 부분 작업 체인(SSC)이라는 상태 전이 표현 방식을 제안합니다. 시연은 구조화된 부분 작업 템플릿(SST) 항목의 연속입니다. 각 SST는 핵심 동작 구성 요소(주어, 술어, 목적어), 유연한 조건(공간적 또는 도구적 구절과 같은 부사 수식어), 팔 동작과 분리된 기본 동작 필드 및 후행 상태 장면 그래프를 저장합니다. 이 형식을 기반으로 SSC는 세 가지 시각-언어 지원 기능을 제공합니다: SST를 자연어로 렌더링, 조립된 체인이 네 가지 상태 전이 규칙에 부합하는지 확인, 그리고 질의 해결 메커니즘을 통해 누락된 필드를 완성합니다. 우리는 논리 검증 및 내용 완성을 위해 BEHAVIOR-1K 데이터셋(50개의 작업, 각 작업당 3개의 에피소드, 2,357개의 주석이 달린 동작 셀)에 이 파이프라인을 적용하고, 13개의 최첨단 시각-언어 모델을 검증기로 사용하여 평가하고, 라벨링 이상 현상을 보고합니다.

Original Abstract

Subtask labels decompose a long-horizon manipulation demonstration into shorter semantic segments for policy training and evaluation. Natural language descriptions are easy to read, but their linguistic variability makes automatic verification difficult. Rigid template formats, such as BEHAVIOR-1K's skill_annotation, are linguistically over-segmented, hindering both readability and annotation consistency. We propose the Structured Subtask Chain (SSC), a state-transition representation that bridges these extremes. A demonstration is a sequence of Structured Subtask Template (SST) entries. Each SST stores core action components (subject, predicate, object), flexible conditions (adverbial modifiers such as spatial or instrumental phrases), a base-motion field separate from arm actions, and an after-state scene graph. Built on this format, SSC supports three vision-language assisted functions: rendering SSTs as natural language, checking the assembled chain against four state-transition rules, and completing underspecified fields through a query resolution cascade. We instantiate the pipeline on BEHAVIOR-1K (50 tasks, 3 episodes per task, 2,357 annotated action cells) for logic verification and content completion, evaluating 13 selected state-of-the-art VL models as candidate verifiers and reporting labelling anomalies.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!