Provenance 기반 입력 그래디언트 가이드 기반 합성 데이터 학습
Learning from Synthetic Data via Provenance-Based Input Gradient Guidance
합성 데이터를 활용한 학습 방법은 훈련 데이터의 다양성을 높이고 데이터 수집 비용을 절감하여 모델의 판별력을 향상시키는 효과적인 접근 방식으로 주목받고 있습니다. 그러나, 기존의 많은 방법들은 훈련 샘플의 다양성을 통해 간접적으로만 강건성을 향상시키며, 모델에게 입력 공간의 어떤 영역이 실제로 판별에 기여하는지 명시적으로 가르치지 않습니다. 따라서, 모델은 합성 과정에서 발생하는 편향 및 인공물로 인한 허위 상관관계를 학습할 수 있습니다. 이러한 한계에 착안하여, 본 논문에서는 훈련 데이터 합성 과정에서 얻은 provenance 정보를 활용하는 학습 프레임워크를 제안합니다. provenance 정보는 입력 공간의 각 영역이 대상 객체에서 비롯되었는지 여부를 나타내는 보조적인 감독 신호로 활용되어, 대상 영역에 집중된 표현 학습을 촉진합니다. 구체적으로, 입력 그래디언트를 합성 과정 중 대상 및 비대상 영역에 대한 정보에 따라 분해하고, 비대상 영역에서의 그래디언트를 억제하는 입력 그래디언트 가이드를 도입합니다. 이를 통해 모델이 비대상 영역에 의존하는 것을 억제하고, 대상 영역에 대한 판별적인 표현 학습을 직접적으로 촉진합니다. 제안하는 방법은 약하게 감독되는 객체 위치 추정, 시공간 행동 위치 추정, 이미지 분류 등 다양한 작업 및 모달리티에서 효과적이고 일반화 가능하다는 것을 실험을 통해 입증합니다.
Learning methods using synthetic data have attracted attention as an effective approach for increasing the diversity of training data while reducing collection costs, thereby improving the robustness of model discrimination. However, many existing methods improve robustness only indirectly through the diversification of training samples and do not explicitly teach the model which regions in the input space truly contribute to discrimination; consequently, the model may learn spurious correlations caused by synthesis biases and artifacts. Motivated by this limitation, this paper proposes a learning framework that uses provenance information obtained during the training data synthesis process, indicating whether each region in the input space originates from the target object, as an auxiliary supervisory signal to promote the acquisition of representations focused on target regions. Specifically, input gradients are decomposed based on information about target and non-target regions during synthesis, and input gradient guidance is introduced to suppress gradients over non-target regions. This suppresses the model's reliance on non-target regions and directly promotes the learning of discriminative representations for target regions. Experiments demonstrate the effectiveness and generality of the proposed method across multiple tasks and modalities, including weakly supervised object localization, spatio-temporal action localization, and image classification.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.