자율 분석 에이전트의 혁신-잔차 감사: 위치 추적, 검출 한계, 오류 제어 및 식별 가능성
Innovation-Residual Auditing of Autonomous Analysis Agents: Localization, Detection Limits, Error Control, and Identifiability
현재 자율 에이전트는 데이터 분석 전체 과정을 수행하며, 집단 선택, 테이블 병합, 모델 적합 등 다양한 작업을 단계별 감독 없이 처리합니다. 이러한 분석 결과가 잘못된 경우, 어떤 작업이 원인이 되었는지 파악해야 합니다. 최근에는 레이블링된 오류 정보 없이, 이미 알려진 정확한 분석 결과를 학습하여 예측 모델에서 벗어나는 작업을 식별하는 방법이 제시되었습니다. 하지만 이러한 감사 방식의 신뢰성은 아직 연구되지 않았습니다. 본 논문은 이러한 문제에 대한 분석을 제공합니다. 점수의 선택은 오류를 아예 특정할 수 있는지 여부를 결정합니다. 각 작업이 이전 작업에 비해 얼마나 예상 밖인지에 따라 점수를 매기면, 단순히 이전 오류를 상속하는 작업은 올바른 작업과 구별되지 않으므로 하나의 오류가 하나의 경고 신호를 발생시킵니다. 반면, 의도된 분석 과정을 더 길게 재구성하여 계산된 점수는 단일 오류를 여러 작업에 분산시킬 수 있습니다. 우리는 이러한 분산 정도를 정량화하고, 점진적으로 오류가 누적될 때 비교 길이를 어떻게 선택해야 하는지 제시합니다. 또한, 개별 감사 분석 내에서 잘못 식별되는 작업의 비율을 제어하는 방법을 제공하며, 이 방법은 적합된 모델이 정확할 필요 없이, 정확한 분석 결과들이 교환 가능하기만 하면 됩니다. 더 나아가, 모델이 불완전하거나, 분석 내용에 따라 검토 대상이 선택되는 경우 보장 수준이 어떻게 약화되는지 정량적으로 분석합니다. 마지막으로, 이러한 감사 방식이 보고할 수 있는 한계를 제시합니다. 특정 크기 이하의 오류는 일반적인 분석 결과의 변동성과 구별할 수 없으므로 전혀 식별할 수 없습니다. 이 한계는 수집된 정확한 분석 결과의 수가 증가함에 따라 매우 느리게 감소하므로, 현재 사용되는 데이터 규모에서 백 배 증가해도 2% 미만으로 감소합니다. 따라서 학습 데이터의 양보다는 표현의 차원이 제약 요인이 됩니다.
Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step supervision. When such an analysis turns out to be wrong, someone must determine which operation caused it. A recent approach does this without any labelled mistakes, learning instead from analyses known to be sound and flagging operations that depart from what that model predicts; how reliable such audits are has not been studied. This paper supplies that analysis. The choice of score determines whether an error can be localized at all. If each operation is scored by how surprising it is given the operation immediately preceding it, then operations that merely inherit an earlier error are indistinguishable from correct ones, so one mistake produces one flag; scores computed against a longer reconstruction of the intended analysis instead spread a single mistake across many operations. We quantify how far they spread, and how to choose the comparison length when an error accumulates gradually rather than at once. We then give procedures that control the proportion of falsely flagged operations within a single audited analysis, requiring only that sound analyses be exchangeable rather than that the fitted model be correct, and we quantify how much the guarantees weaken when the model is imperfect or when the analysis was selected for review in a way that depends on its content. Finally we establish a limit on what any such audit can report: errors below a certain magnitude cannot be attributed at all, being indistinguishable from ordinary variation among sound analyses. This limit falls so slowly as more sound analyses are collected that at the representation sizes now in use a hundredfold increase reduces it by under two percent, so the dimension of the representation rather than the volume of training data is the binding constraint.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.