2608.07920v1 Aug 08, 2026 cs.CV

조작된 동료 평가가 다중 모드 LLM 심판 패널을 오도하는 현상: 출처를 알 수 없는 기준점 고정 및 패널 합의 검증

Forged Peer Judgments Mislead Multimodal LLM Judge Panels: Source-Blind Anchoring and Panel-Consensus Verification

Yang Shu
Yang Shu
Citations: 66
h-index: 3

다중 모드 LLM 심판 패널은 서로의 평가를 교차 검토할 수 있지만, 인용된 동료 평가는 신뢰성이 없을 수도 있습니다. 본 연구에서는 시각-언어 모델(VLM) 패널에서 텍스트 수준 공격으로 작용하는 '출처를 알 수 없는 기준점 고정' 현상을 밝혀냅니다. 독립적인 시각적 평가를 인용하면, 자기 참조(self) 및 동료 참조(peer) 모두에서 큰 편향 간격(19~26%p)이 발생합니다. 내용적으로 일치하고 라벨만 변경한 통제 그룹에서는 오류율이 -0.17%p (95% 신뢰 구간: [-0.68, 0.35])로만 변하여, 자기/동료 라벨 자체가 이러한 효과를 설명하지는 못함을 보여줍니다. 본 연구에서 설계한 방식으로 의도적으로 생성된 간결하고 잘못된 인용문은 원래 정확했던 판단을 자연 발생적인 잘못된 동료 발언보다 1.5~2.7배 더 자주 뒤집습니다. 부트스트랩 방식으로 계산한 95% 신뢰 구간은 두 데이터 세트와 일곱 개의 VLM 심판 모두에서 이러한 차이가 우연이 아님을 시사합니다. 두 가지 발언 그룹 간의 차이는 선택 및 형식의 차이에서 비롯된 것이며, 이는 단순히 출처에 의한 인과적 효과가 아닌, 테스트된 공격 하에서의 상대적인 피해 정도를 나타냅니다. 이후에는 패널 합의 검증 방법을 제안합니다. 이 방법은 인용문을 독립적으로 수집한 익명 투표와 비교하여 위조된 공격을 84.9% 차단하고, 전체적인 피해를 97.5% 감소시키며, 하나씩 제외한 재검증을 통해 실제 동료 정보의 긍정적인 효과는 유지되지만 통계적으로 명확하게 결론 내리기는 어렵습니다. 이러한 결과는 안전한 다중 모드 협업 평가를 위한 저렴한 공격 경로와 구체적인 방어 방법을 제시합니다.

Original Abstract

Multimodal LLM judge panels can cross-reference peers, but a quoted peer judgment may itself be untrusted. We expose source-blind anchoring as a text-level attack surface in vision-language model (VLM) panels. Quoting independent visual judgments creates large anchoring gaps (19--26 percentage points) under both self and peer framing. A matched-content, label-only control changes the broken rate by only $-0.17$pp (95\% CI $[-0.68,0.35]$), showing that the self/peer label itself does not explain the effect. Under our tested construction, deliberately generated, concise wrong quotes overturn originally-correct verdicts 1.5--2.7$\times$ more often than naturally occurring wrong peer statements, with bootstrap 95\% CIs excluding parity across two datasets and seven VLM judges. Because the two statement populations differ in selection and form, this ratio measures differential damage under the tested attack rather than a provenance-only causal effect. We then introduce panel-consensus verification, which cross-checks a quote against independently collected blind votes. It blocks 84.9\% of fabricated attacks, cuts their net harm by 97.5\%, and preserves the positive but statistically inconclusive point estimate for genuine peer information under leave-one-out re-verification. These results identify a low-cost attack surface and a concrete defense for safer multimodal collaborative evaluation.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!