지원 연산 분해: 제어된 개입 하에서의 고정 시각 인코더의 조합적 읽기
Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions
고정된 시각 인코더에 대한 조합 분석은 무엇이 변경되었는지, 그리고 어디가 변경되었는지를 파악해야 합니다. 기존의 팩터 프로브는 이러한 축들을 개별적으로 평가하지만, 동일한 예측 슬롯을 재사용하는 여러 연산을 보상할 수 있으며, 이를 우리는 '연산 세탁'이라고 부릅니다. 우리는 지원-연산 그리드에 대한 주입적으로 정렬된 leave-one-cell-out 프로토콜과 SO-OPF라는 읽기 방식을 도입합니다. SO-OPF는 셀 에너지를 지원의 중요도와 경쟁적인 연산의 사후 확률로 분해합니다. 이러한 공식은 집계 점수가 혼동시키는 두 가지 질문을 분리합니다. 즉, 그리드가 알려진 경우 캐리어가 생략된 바인딩을 구성하는지 여부와 평탄한 셀 레이블로부터 해당 그리드를 복구할 수 있는지 여부를 판단하는 것입니다. 고정된 DINOv3 특징을 사용했을 때, 알려진 팩토리얼 할당은 Shapes3D-Extended 데이터셋에서 0.874의 주입 정확도를, 전역적으로 이미지 분리된 COCO 데이터셋에서 0.799의 정확도를 달성했습니다. 평탄한 레이블로부터 할당을 학습했을 때 각각 0.769와 0.762의 정확도를 얻었습니다. Shapes3D 데이터셋에서 축 일치에 대한 감독 학습 하에서, 분해된 캐리어는 밀집된 캐리어보다 학습된 할당 정확도를 0.653에서 0.841로 향상시키고 '연산 세탁' 문제를 해결합니다. SigLIP2는 COCO 데이터셋에 대해 유사한 분리를 재현합니다. 재구축된 MuJoCo 환경에서는 경계가 드러납니다. DINOv3를 사용했을 때 학습된 할당 정확도는 0.569이고, SigLIP2를 사용할 때는 0.484로 떨어지며, 상당한 슬롯 붕괴가 발생합니다. 따라서 분해된 읽기 방식과 주입적 평가 방법은 두 가지 환경에서 생략된 바인딩을 복구하는 동시에 특정 렌더링 엔진에 따른 문제점을 드러내지만, 평탄한 레이블로부터 보편적인 복구를 가능하게 하지 않습니다.
Compositional analysis of frozen vision encoders should determine both what changed and where it changed. Standard factor probes score these axes separately, however, and can reward multiple operations that reuse the same predicted slot. We call this failure operation laundering. We introduce an injectively aligned leave-one-cell-out protocol over support x operation grids and SO-OPF, a readout that factors cell energy into support salience and a competitive operation posterior. This formulation separates two questions that aggregate scores conflate: whether the carrier composes held-out bindings when the grid is known, and whether that grid can be recovered from flat cell labels. With frozen DINOv3 features, known factorial assignment reaches 0.874 injective accuracy on Shapes3D-Extended and 0.799 on globally image-disjoint COCO; learning the assignment from flat labels reaches 0.769 and 0.762, respectively. Under matched-axis-aware supervision on Shapes3D, the factored carrier improves learned-assignment accuracy from 0.653 to 0.841 over a dense carrier and eliminates its laundering gap. SigLIP2 replicates the COCO separation. A rebuilt MuJoCo substrate exposes a boundary: learned-assignment accuracy is 0.569 with DINOv3 and 0.484 with SigLIP2, with substantial slot collapse. Thus factored readout and injective evaluation recover held-out bindings on two substrates while exposing, rather than hiding, a renderer-specific failure boundary; they do not establish universal recovery from flat labels.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.