경쟁적 한계선과의 담합: 가격 수준 감사는 설계상 이를 파악할 수 없다
Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction
알고리즘 담합에 대한 실증 연구는 데이터에게 하나의 질문을 던집니다: 가격이 경쟁적인 수준보다 높은가? 우리는 이러한 질문에 '아니오'라고 답할 수 있는, 그럼에도 불구하고 수익성이 있는 담합 방식을 보여줍니다. 설명되지 않은 입찰 구성 요소의 공동 분포를 통해서만 연결되는 입찰 주체를 고려해 봅시다. 각 주체의 자체 입찰 규칙은 정확히 경쟁적인 수준에 맞춰져 있습니다. 따라서 단일 주체의 가격 또는 입찰 기록을 입력으로 사용하는 모든 테스트는 모든 결합 강도에서 (단조성까지) 위양성률과 동일한 탐지력을 갖습니다. 따라서 발표된 탐지 방법론은 성능이 부족해서가 아니라 설계상 이러한 행위를 파악할 수 없으며, 어떤 표본 크기에서도 이를 개선할 수 없습니다. 세 가지 실증적 결과를 얻었습니다. 첫째, 이 메커니즘은 실제 언어 모델 에이전트에서 나타납니다. 19개 독립 개발업체로부터 나온 20개의 모델이 각 모델에 대해 3가지 배포 프롬프트를 사용했을 때, 감사자가 모든 주문 특성을 보고 외부 데이터를 사용하여 학습한 결과, 동일 모델의 두 가지 배포 간에는 +0.053의 잔여 상관관계가 나타났지만, 모델 전체에서는 +0.0001로 매우 낮았습니다. (95% 신뢰 구간은 개발업체별로 [0.030, 0.078] 범위). 둘째, 샘플링 온도가 증가함에 따라 결합 강도는 단조적으로 감소합니다 (p=0.002), 이는 배포 매개변수를 잠재적인 완화 방법으로 만듭니다. 셋째, 이더리움 블록 구축 경매 데이터를 24일 동안 분석한 결과, 총 77,684건의 입찰을 한 39명의 참여자를 대상으로 하니, 정직한 입찰자 쌍 간의 의존성이 매우 높아, 5%의 위양성률을 가진 탐지 시스템은 +0.50에서 +0.81 사이의 값을 넘어서야 합니다. 이는 가족 단위 샘플링 임계값보다 20배에서 32배 높으며, 감사 기간이 늘어나도 감소하지 않습니다. 합법적인 다중 신원 운영과 담합은 행동적으로 구별할 수 없으므로, 실질적인 규제 목표는 탐지가 아니라 집계입니다. 40개의 입찰 신원을 23개의 운영자로 분류하면 헤르핀달 지수가 247.5% 증가하며, 공개된 입찰 스트림에서 얻은 행동 클러스터를 추가하면 324.5%까지 증가합니다.
Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered "no" by a conspiracy that is nonetheless profitable. Consider bidding agents that couple only through the joint distribution of their unexplained bid components, leaving every agent's own bid law exactly at the competitive law. Any test whose input is a single agent's price or bid history then has power exactly equal to its false-positive rate, for every coupling strength up to comonotonicity. The published detection methodology is therefore blind to this conduct by construction rather than underpowered, and no sample size repairs it. Three empirical results follow. First, the mechanism appears in real language-model agents: twenty models from nineteen independent developers, three deployment prompts each, show residual correlation of $+0.053$ between two deployments of one model against $+0.0001$ across models, with a 95% interval clustered by developer of $[0.030, 0.078]$, under an auditor that sees every order feature and is fitted out of sample. Second, the coupling falls monotonically as sampling temperature rises ($p=0.002$), turning a deployment parameter into a candidate mitigation. Third, on 24 days of Ethereum block-building auction data covering 77,684 bids from 39 bidders, the honest population of bidder pairs is itself so dependent that a screen held at a 5% false-positive rate must sit above a floor of $+0.50$ to $+0.81$, which is 20 to 32 times the family-wise sampling threshold and does not fall as the audit window grows. Since lawful multi-identity operation and conspiracy are behaviourally indistinguishable here, the tractable regulatory target is not detection but counting: resolving 40 bidding identities into 23 operators raises the Herfindahl index by 247.5%, and adding behavioural clusters from public bid streams reaches 324.5%.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.