환영 참고문헌: 최고 수준 학회에서 살아남은 허구적인 인용
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
대규모 언어 모델은 뒷받침되지 않는 주장을 포함하는 정교한 과학 텍스트를 생성할 수 있으며, 이는 '환각' 현상을 통해 기록에 남게 됩니다. 이러한 위험을 기술적 방법으로 평가하기는 어렵고 종종 전문가의 판단이 필요하지만, 인용문은 보다 감사하기 쉬운 영역입니다. 인용문은 실제 학술 연구 결과로 연결되고 저자 정보가 일치하거나, 그렇지 않으면 오류라는 것을 알 수 있습니다. 저희는 엄격한 기준(존재하지 않는 자료 및 상당한 저자 목록 불일치)에 따라 동료 평가를 거친 논문에서 인용 환각 현상을 측정했습니다. 일반적인 서지 정보의 변경 사항(예: 발표 장소/연도 차이, 출판 상태 업데이트, 경미한 이름 변형)은 명시적으로 제외했습니다. 대규모로 인용문을 감사하기 위해 저희는 여러 서지 자료와 비교하여 참고문헌 항목을 확인하고 해결되지 않은 경우 웹 검색을 통해 재검증하는 검증 파이프라인인 RefChecker를 개발했습니다. 저희는 ICLR, ICML, NeurIPS 및 USENIX Security에 채택된 최종 논문에 RefChecker를 적용했습니다. 허구적인 인용문이 기록에 남아 있습니다. 참고문헌 수준의 오류율은 일반적으로 1% 미만이지만, 학회 발표 자료의 규모가 커서 개별 논문의 오류가 눈에 띄게 나타납니다. 저희 기준에 따르면, 2025년에 NeurIPS 및 USENIX Security 논문 약 20개 중 1개가 최소 두 개의 '환각'으로 보이는 학술 논문과 관련된 인용문을 포함하고 있습니다. 또한 ChatGPT 등장 이후 일부 분야에서 이러한 현상이 증가했으며, 심지어 수상 논문에서도 5개 이상의 오류가 있는 참고문헌 목록이 발견되고 '환각'으로 보일 수 있는 인용문이 존재합니다. 이러한 결과는 동료 평가만으로는 인용의 정확성을 신뢰할 수 없을 정도로 보장하지 못하지만, 감사는 실현 가능하며(특정 학회 규모의 스캔 시 논문당 약 0.04달러). 저희는 RefChecker를 공개 소스 프로젝트로 제공하여 출판 전에 정기적이고 재현 가능한 인용 검증을 수행할 수 있도록 지원합니다 (https://github.com/markrussinovich/refchecker).
Large language models can generate polished scientific text that includes unsupported claims, allowing hallucinations to enter the archival record. Assessing this risk via technical statements is difficult and often requires expert judgment, but citations provide a more auditable surface: a reference either resolves to a real scholarly work with compatible authorship, or it does not. We measure citation hallucination in peer-reviewed proceedings using a conservative definition limited to identity-level failures: non-existent works and substantial author-list mismatches. We explicitly exclude ordinary bibliographic drift (e.g., venue/year differences, publication-status updates, minor name variants). To audit citations at scale, we build RefChecker, a verification pipeline that resolves bibliography entries against multiple bibliographic sources and escalates unresolved cases to web-search re-verification. We apply RefChecker to accepted camera-ready papers from ICLR, ICML, NeurIPS, and USENIX Security. Hallucinated citations have entered the archival record. While reference-level rates are usually below 1%, proceedings are large enough that paper-level failures are visible: in 2025, roughly one in twenty NeurIPS and USENIX Security papers contains at least two likely hallucinated academic-paper-like references under our strict definition. We also observe post-ChatGPT increases in several venues, including a tail of papers with 5+ failures in a single bibliography, and likely hallucinated citations even among award-winning papers. These results suggest peer review alone does not reliably enforce citation integrity, yet auditing is tractable (about 0.04$ per paper in one venue-scale scan). We open-source RefChecker for routine, reproducible citation verification before publication (https://github.com/markrussinovich/refchecker).
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.