2607.27031v1 Jul 29, 2026 cs.LG

로또 티켓은 배포 티켓이 아니다

Lottery Tickets Are Not Deployment Tickets

Bum Jun Kim
Bum Jun Kim
The University of Tokyo
Citations: 271
h-index: 8

과거 연구에서 희소화, 압축 및 로또 티켓이 모델 동작에 미치는 영향에 대한 보고는 상반된 결과를 보였습니다. 일부 연구에서는 긍정적인 효과가 관찰되었지만, 다른 연구에서는 부정적인 효과가 나타났습니다. 또한, 기존 연구에서는 의사 결정 로직이 이미 고정되어 있는 실제 배포 환경을 고려하지 않았습니다. 본 연구는 이러한 상반된 결과들을 실용적인 관점에서 평가하기 위해, 정확도를 맞춘 로또 티켓 또는 다른 희소 모델이 기존의 밀집 모델을 대체할 때, 하위 의사 결정 로직을 재구성하지 않고도 가능한지 여부를 배포 수준에서 조사합니다. 이를 위해, 보정(calibration), OOD 응답, 클래스별 신뢰성, 표현 및 하위 정책 결정 등과 같이 배포와 관련된 다양한 동작들을 프로토콜별로 분석하고, 정확도를 제외한 동작 변화를 '행동적 호환성 거리'라는 지표로 요약합니다. 광범위한 실험 결과, 희소 모델은 반복적으로 기존 밀집 모델의 정확도를 회복하지만, 여전히 행동적으로는 차이를 보입니다. 특히 특정 조건에서 로또 티켓은 낮은 부정확도(corruption accuracy)를 보이기도 합니다. 고정된 임계값을 사용하는 정책 진단 환경에서는, 로또 티켓으로 대체할 때 수락-검토 결정의 7%에서 10%가 변경됩니다. 이러한 변화는 '간편하게 교체 가능'하도록 설계된 시스템이 해결하고자 하는 문제점, 즉 하위 의사 결정 로직을 재구성하고 재검증해야 한다는 부담을 야기합니다. 본 연구 결과는 정확도만을 기준으로 모델의 안전성을 평가하는 데 한계가 있음을 보여줍니다. 기존 모델과의 호환성을 확보하는 것은, 변화를 단순히 희소성으로 돌리거나 모든 측정된 차이를 유해하다고 간주하는 것과는 다른 문제입니다. 이론적 분석을 통해 이러한 결과를 설명하며, 정확도가 완벽하게 일치하더라도 고정된 임계값 기반의 의사 결정 변경은 발생할 수 있으며, 운영 경계 근처에서 발생하는 작은 신뢰도 변화가 상당한 의사 결정 변화를 초래할 수 있다는 점을 강조합니다.

Original Abstract

Reports on how sparsification, compression, and lottery tickets change model behavior have been mixed in the prior literature, with beneficial effects observed in some studies and adverse effects in others. Moreover, prior work has not considered actual deployment conditions, where decision logic is already fixed for the incumbent. To assess these mixed findings from a practical standpoint, we study the production-replacement question at the deployment level, namely whether an accuracy-matched lottery ticket or another sparse challenger can replace an incumbent dense model without reconfiguring downstream decision logic. We therefore audit a broad, protocol-specific panel of deployment-relevant behaviors spanning calibration, OOD response, class-level reliability, representations, and downstream policy decisions, and summarize clean-accuracy-excluded deviations with a behavioral-compatibility distance. Across extensive experiments, sparse candidates repeatedly recover dense-reference accuracy yet remain behaviorally different; in several study-band-matched settings, LTs also show lower corruption accuracy. In small-gap settings with fixed-threshold policy diagnostics, lottery-ticket replacement changes 7% to 10% of accept--review decisions. This churn creates precisely the burden that drop-in replacement is meant to avoid: reconfiguring and revalidating downstream decision logic. These findings establish the limits of clean-accuracy certification: Establishing compatibility with a fixed incumbent is distinct from attributing churn uniquely to sparsity or treating every measured deviation as harmful. Our theory explains the routing result: Even exact pointwise top-1 agreement cannot bound fixed-threshold decision changes, and small confidence shifts near the operating boundary can generate first-order routing churn.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!