2607.24884v2 Jul 27, 2026 cs.SE

검색 증강 코드 생성에서의 불확실성: '무엇을 검색할 것인가'를 넘어

Beyond "What to Retrieve": Uncertainty in Retrieval-Augmented Code Generation

Li Zhang
Li Zhang
Citations: 42
h-index: 3
Chandan Kumar Sah
Chandan Kumar Sah
Beihang University
Citations: 80
h-index: 6
Xiaoli Lian
Xiaoli Lian
Citations: 38
h-index: 4

저장소 수준의 코드 생성은 관련성, 호환성 및 완전성이 본질적으로 불확실한 이종 증거에 의존합니다. 유사 코드를 활용한 예시, 저장소 컨텍스트 및 프로젝트별 API는 상호 보완적인 정보를 제공할 수 있지만, 동시에 노이즈가 있거나 중복되거나 충돌하는 신호를 유발할 수도 있습니다. 기존의 검색 증강 방식은 주로 검색 관련성을 최적화하는 데 집중하며, 검색된 증거에 내재된 불확실성이 후속 생성 과정에 미치는 영향을 명시적으로 모델링하지 않습니다. 본 논문에서는 소스별 불확실성을 추정하고, 이를 활용하여 이종 증거를 필터링하고 순위를 매기며, 생성, 검증 및 수정을 안내하는 불확실성 인지 프레임워크인 OpenCoder를 소개합니다. API 지식, 저장소 컨텍스트 및 유사 코드 증거에 대한 요인 분석 결과, 보편적인 가중치 합산 순위는 존재하지 않으며, 대신 상당한 상호 작용이 발생하며, 이는 함께 사용되는 증거와 LLM 백엔드에 따라 달라집니다. 확장된 32개 작업의 RepoExec-inline 평가에서 OpenCoder는 GPT가 선택한 출력의 정확도를 Baseline RAG 모델보다 56.25%에서 78.13%로 향상시켰습니다. 그러나 검증 및 수정을 수행하는 제어 그룹과 동일한 성능을 보였으며, Gemini를 사용했을 때의 개선 효과는 통계적으로 유의미하지 않았습니다. 또한, 목표 인지 API 정제 기술은 API 집합 검색 성능을 크게 향상시켰습니다. 이러한 결과는 불확실성을 저장소 수준의 검색, 검증 및 수정을 위한 실행 가능한 제어 신호로 간주할 수 있음을 시사합니다.

Original Abstract

Repository-level code generation relies on heterogeneous evidence whose relevance, compatibility, and completeness are inherently uncertain. Similar-code examples, repository context, and project-specific APIs may provide complementary information, but can also introduce noisy, redundant, or conflicting signals. Existing retrieval-augmented approaches primarily optimize retrieval relevance without explicitly modeling how uncertainty in retrieved evidence affects downstream generation. We introduce OpenCoder, an uncertainty-aware framework that estimates source-specific uncertainty, uses it to filter and rank heterogeneous evidence, and guides generation, verification, and repair. A factorial analysis over API knowledge, repository context, and similar-code evidence reveals no universal additive source ranking; instead, significant cross-source interactions depend on the accompanying evidence and LLM backend. On an expanded 32-task RepoExec-inline evaluation, OpenCoder improves GPT selected-output correctness over Baseline RAG from 56.25\% to 78.13\%. However, it matches a verification-and-repair control, and the corresponding Gemini improvement is not statistically supported, indicating backend-dependent benefits. Target-aware API refinement also substantially improves API-set retrieval. These findings support treating uncertainty as an actionable control signal for repository-level retrieval, verification, and repair.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!