2608.03216v1 Aug 04, 2026 cs.CV

iFAN: 추론 인지 학습을 통한 단순 마스크 변환기

iFAN: Inference-Aware Learning for Plain Mask Transformers

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Haoyang Tong
Haoyang Tong
Citations: 71
h-index: 5
Wenxiao Fan
Wenxiao Fan
Citations: 1
h-index: 1
Lichen Ma
Lichen Ma
Citations: 29
h-index: 2
Jingling Fu
Jingling Fu
Citations: 17
h-index: 2
Lu Liu
Lu Liu
Citations: 55
h-index: 3
Junshi Huang
Junshi Huang
Citations: 19
h-index: 2

쿼리 기반 마스크 변환기는 최종 레이어의 쿼리 예측 간의 픽셀 단위 경쟁을 통해 분할 결과를 조립하지만, 이 추론 과정은 학습 중에 명시적으로 최적화되지 않습니다. 우리는 두 가지 주요 불일치를 발견했습니다. 가장 높은 확률-마스크 점수를 가진 쿼리가 반드시 가장 정확한 마스크를 생성하는 것은 아니며, 최종 레이어 디코딩이 중간 레이어에서 생성된 우수한 예측을 버릴 수 있습니다. 이러한 문제를 해결하기 위해, 우리는 단순 마스크 변환기를 위한 일반적인 학습 프레임워크인 추론 인지 학습 (iFAN)을 제안합니다. iFAN은 쿼리 경쟁을 예측된 마스크 품질에 맞추고, 높은 신뢰도를 가지지만 부정확한 경쟁자를 억제하는 조정된 확률-마스크 순위 (APMR)를 도입합니다. 또한, 더 강력한 중간 레이어 예측을 최종 레이어로 전달하기 위해 크로스 레이어 자기 증류 (CLSD)를 사용합니다. 순위 및 증류 목표는 학습 과정에서만 적용되며, 추론은 효율적인 최종 레이어 디코딩을 유지합니다. COCO, ADE20K 및 Cityscapes 데이터셋에 대한 실험 결과, iFAN은 다양한 아키텍처, 백본 크기 및 입력 해상도에서 판옵틱, 인스턴스 및 의미 분할 성능을 일관되게 향상시킵니다. 전반적으로, iFAN은 평균 1.20 PQ, 1.30 AP 및 0.63 mIoU의 성능 향상을 가져오며, 추가적인 파라미터, FLOPs 및 추론 지연 시간은 미미합니다.

Original Abstract

Query-based mask transformers assemble segmentation outputs through pixel-wise competition among query predictions of the final layer, yet this inference process is not explicitly optimized during training. We identify two key mismatches: the query with the highest probability-mask score does not necessarily produce the most accurate mask, and final-layer decoding may discard superior predictions from intermediate layers. To address these issues, we propose Inference-Aware Learning (iFAN), a general training framework for plain mask transformers. iFAN introduces Adjusted Probability-Mask Ranking (APMR), which aligns query competition with predicted mask quality and suppresses high-confidence but inaccurate competitors. We further employ Cross-Layer Self-Distillation (CLSD) to transfer stronger intermediate predictions to the final layer. The ranking and distillation objectives are training-only, while inference retains efficient final-layer decoding. Experiments on COCO, ADE20K, and Cityscapes demonstrate consistent improvements across panoptic, instance, and semantic segmentation, as well as across different architectures, backbone scales, and input resolutions. Overall, iFAN improves performance by an average of 1.20 PQ, 1.30 AP, and 0.63 mIoU, with negligible additional parameters, FLOPs and inference latency.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!