대체 정보 사용이 개인 식별 정보(PHI) 탐지 가능성을 유지하는가: 다중 검출기 동등성 연구
Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study
구조 보존 식별 제거는 보호된 건강 정보(PHI)를 실제적인 동일 유형의 대체 정보로 대체합니다. 예를 들어, "Anna S."는 "Maria S."로 변경되며, [NAME]과 같은 일반적인 표현으로 대체되지 않습니다. 이를 통해 임상 텍스트의 흐름이 유지되고 후속 도구가 계속 작동할 수 있습니다. 그러나 이러한 방법이 유효하려면 대체 정보 자체가 해당 도구가 의존하는 신호를 손상시키지 않아야 합니다. 본 연구에서는 다음과 같은 구체적이고 검증 가능한 질문을 던집니다: 식별 제거 과정에서 실제로 마스킹되는 영역에서, 후속 PHI 탐지기가 여전히 대체 정보를 찾을 수 있는가? 본 연구는 (i) 커버리지와 유용성을 분리하여 유용성 평가를 마스킹된 영역에만 적용하고, (ii) 표본 크기(57,000개의 쌍으로 구성됨)에서 유의미하지 않은 귀무 가설 검정 대신 동등성 검정(TOST)을 사용하며, (iii) 수정 가능한 생성기 결함과 고유한 탐지기 한계를 구분하는 대체 정보 실패 유형론을 구축하는 이중화된 다중 탐지기 평가 프로토콜을 소개합니다. 11개의 탐지기, 7개의 벤치마크 및 7개 언어(총 1,750개 문서)에 대한 분석 결과, 마스킹된 영역에서의 재현율은 76.1%에서 74.9%로 감소했습니다. 동등성 검정을 통해 이러한 변화는 +/- 2 포인트 이내의 범위 내에서 통계적으로 0과 동일한 것으로 나타났으며(p ~ 3e-9), 탐지기 순위는 유지되었습니다. 잔여 손실은 PHI 탐지기가 성능이 저하되었음을 의미하지 않으며, 잘못 구성되거나 데이터 분포와 일치하지 않는 대체 정보로 인해 발생하는 현상입니다(예: "Chicago"가 "Illino"로, "Cedars-Sinai"가 "Vidant"로 변경됨). 삭제 기준 및 오픈 소스 대체 정보 기본값을 통해 이러한 효과는 특정 도구의 문제가 아니라 잘 구성된 대체 정보 자체의 특성임을 알 수 있습니다. 본 연구에서는 평가 데이터 세트, 점수 계산 코드 및 대화형 대시보드(https://custodianai.pages.dev)를 공개하여 이 프로토콜을 사용하여 모든 구조 보존 변환을 감사할 수 있도록 했습니다.
Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S." becomes "Maria S.", not [NAME] -- so that clinical text stays fluent and downstream tools keep working. But this only helps if the substitution does not itself corrupt the signal those tools rely on. We ask a narrow, testable question: on the spans a de-identifier actually masks, can downstream PHI detectors still find the surrogate? We introduce a paired, multi-detector evaluation protocol that (i) scores utility only on masked spans, decoupling coverage from utility; (ii) uses equivalence testing (TOST) rather than null-hypothesis significance testing, which is uninformative at our sample size (57k paired spans); and (iii) builds a surrogate-failure typology separating fixable generator defects from intrinsic detector limits. Across 11 detectors, 7 benchmarks, and 7 languages (1,750 documents), recall on masked spans moves from 76.1% to 74.9% -- a change our equivalence test shows is statistically equivalent to zero within a +/-2-point margin (p ~ 3e-9), with detector ranking preserved. The residual loss does not reflect detectors getting worse at PHI: it concentrates in malformed and out-of-distribution surrogates (truncation Chicago -> Illino, salience loss Cedars-Sinai -> Vidant). A redaction floor and an open-source surrogate baseline indicate the effect is a property of well-formed substitution, not of one tool. We release the evaluation subsets, scoring code, and an interactive dashboard at https://custodianai.pages.dev so the protocol can audit any structure-preserving transform.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.