2606.11804v1 Jun 10, 2026 cs.AI

신뢰할 수 있는 인공지능을 향하여: 연속 데이터 요약에 대한 다중 목표 적대적 공격 및 견고한 방어

Toward Trustworthy AI: Multi-Target Adversarial Attacks and Robust Defenses for Continuous Data Summarization

Jason Xue
Jason Xue
Citations: 23
h-index: 2
Yanan Cai
Yanan Cai
Citations: 7
h-index: 2
Shuchao Pang
Shuchao Pang
Citations: 134
h-index: 7
Yuefang Lian
Yuefang Lian
Citations: 5
h-index: 1
Longkun Guo
Longkun Guo
Citations: 9
h-index: 2
Zhigang Lu
Zhigang Lu
Citations: 70
h-index: 4
Dachuan Xu
Dachuan Xu
Citations: 57
h-index: 2
Zhongrui Zhao
Zhongrui Zhao
Citations: 3
h-index: 1

신뢰할 수 있는 인공지능은 견고한 예측 모델뿐만 아니라 신뢰성 있는 데이터 처리 파이프라인을 요구합니다. 상위 단계 구성 요소인 데이터 요약은 어떤 정보가 유지되고 후속 학습 또는 의사 결정 모듈로 전달되는지를 결정합니다. 따라서 요약 과정에 대한 적대적 공격은 상위 단계에서 신뢰할 수 있는 인공지능을 위협할 수 있습니다. 이러한 공격은 선택된 요약을 변경하고, 대표성을 감소시키며, 후속 학습 작업의 유용성을 더욱 저하시킬 수 있습니다. 본 논문에서는 DR-부분집합 최적화를 통해 유사성 수준의 교란이 발생하는 연속 데이터 요약에 대한 적대적 공격을 연구합니다. 우리는 다중 해상도 이미지 요약 목표가 음수가 아닌 부분집합 함수를 이용한 다선형 확장으로 표현될 수 있으며, $m$-약강 단조성을 만족하는 DR-부분집합 특성을 갖는다는 것을 보여줍니다. 그런 다음, 우리는 여러 대상 요약 모델의 성능을 저하시키기 위해 유사성 구조에 대한 허용 가능한 하나의 교란을 최적화하는 최소-최대 문제로 다중 목표 공격 생성을 공식화합니다. 이러한 교란을 완화하기 위해, 우리는 다양한 유형의 공격에 대한 견고한 방어를 정규화된 최대-최소 문제로 공식화합니다. 두 문제 모두에 대해 이론적인 보장을 갖는 근사 알고리즘을 개발했습니다. 실제 데이터와 제어된 클러스터링 벤치마크에서의 실험 결과, 제안된 공격은 대표적인 저가에서 중간 가격 범위에서 효과적이며, 후속 작업의 성능 손실을 유발할 수 있음을 보여줍니다. 제안된 방어는 구조화된 환경에서 견고성-완화 균형을 개선하며, 실제 데이터에 대한 견고한 보호의 파라미터 민감성을 드러냅니다.

Original Abstract

Trustworthy AI requires reliable data-processing pipelines, not only robust downstream predictive models. As an upstream component, data summarization determines which information is retained and passed to subsequent learning or decision modules. Therefore, adversarial perturbations to the summarization process can compromise trustworthy AI in an upstream manner: they may alter the selected summary, reduce its representativeness, and further degrade the utility of subsequent learning tasks. In this paper, we study adversarial attacks on continuous data summarization under similarity-level perturbations through DR-submodular optimization. We show that a class of multi-resolution image summarization objectives can be formulated as multilinear extensions of non-negative submodular set functions and satisfy DR-submodularity with $m$-weak monotonicity. We then formulate multi-target attack generation as a min-max problem, where one admissible perturbation of the similarity structure is optimized to degrade multiple target summarization models. To mitigate such perturbations, we formulate robust defense against mixed attack types as a regularized max-min problem. For both problems, we develop approximation algorithms with theoretical guarantees. Experiments on real-data and controlled clustered benchmarks show that the proposed attack is effective in representative low-to-moderate budget regimes and can induce downstream task-performance loss. The proposed defense improves the robustness--mitigation trade-off in structured settings, while also revealing the parameter sensitivity of robust protection on real data.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!