2608.04783v1 Aug 05, 2026 cs.SE

RepoProbe: 체크리스트를 활용한 아키텍처 기반 저장소 이해 능력 평가

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists

Zhichao Hu
Zhichao Hu
Citations: 50
h-index: 5
Yuhong Liu
Yuhong Liu
Citations: 318
h-index: 4
Richeng Xuan
Richeng Xuan
Citations: 108
h-index: 4
Alyssa Wu
Alyssa Wu
Citations: 0
h-index: 0
Ji Luo
Ji Luo
Citations: 29
h-index: 2
Zhen Qin
Zhen Qin
Citations: 10
h-index: 2
Yue Yang
Yue Yang
Citations: 0
h-index: 0

대규모 언어 모델(LLM)이 소프트웨어 엔지니어링에 통합되면서, 기능 수준의 코드 생성에서 저장소 전체 규모의 지원으로 초점이 이동하고 있습니다. 그러나 기존 벤치마크는 주로 GitHub Issues의 버그 보고서를 활용하는데, 이는 모델들이 오류 로그 패턴 매칭을 통해 진정한 이해를 회피할 수 있는 가능성을 제공합니다. 이러한 불일치는 Edit Bias(모델이 기존 저장소 아키텍처를 이해하기 전에 코드를 수정하는 경향)를 과소평가합니다. 또한, 현재 LLM 평가에 사용되는 단일 점수 방식은 높은 변동성과 낮은 해석력을 가지고 있습니다. 본 연구에서는 GitHub Discussions를 활용하여 개방형 질의응답을 통해 저장소 수준의 코드 이해 능력을 평가하는 새로운 벤치마크인 RepoProbe를 소개합니다. RepoProbe는 결함 보고서가 아닌, 개방형 아키텍처 관련 질문에 중점을 둡니다. 엄격한 평가를 위해, 우리는 답변을 원자적이고 검증 가능한 사실로 분해하여 주관적인 평점을 객관적인 검증으로 대체하는 체크리스트 기반 검증 프로토콜을 제안합니다. 최첨단(SOTA) LLM에 대한 우리의 평가는 높은 명확성과 증거 기반 기술적 정확성 사이의 지속적인 격차를 보여줍니다. 또한, 모델이 아키텍처 분석보다 코드 생성을 우선시하는 Edit Bias 현상이 실제로 얼마나 흔한지를 정량적으로 확인합니다. 마지막으로, 우리의 검증 프로토콜은 기존의 단일 점수 방식 평가에 비해 평가 신뢰도를 크게 향상시키는 것을 입증했습니다.

Original Abstract

The integration of Large Language Models (LLMs) into software engineering has shifted the focus from function-level generation to repository-scale assistance. However, existing benchmarks largely rely on bug reports from GitHub Issues, which often allow models to bypass genuine understanding via pattern matching on error logs. This misalignment under-measures Edit Bias, which refers to premature generation, where models prematurely propose code modifications instead of understanding the existing repository architecture. Furthermore, current LLM-as-a-Judge scalar scoring suffers from high variance and low interpretability. This work introduces RepoProbe, a novel benchmark for evaluating repository-level code understanding through open-ended Q&A using GitHub Discussions, which focuses on open-ended architectural inquiries rather than defect reporting. To ensure rigorous evaluation, we propose a Checklist-Based Verification Protocol that decomposes answers into atomic, verifiable facts, thereby replacing subjective ratings with objective verification. Our evaluation of state-of-the-art (SOTA) LLMs reveals a persistent gap between high clarity and evidencegrounded technical correctness. It also quantitatively confirms the prevalence of edit bias, in which models prioritize code generation instead of architectural analysis. Finally, we demonstrate that our verification protocol significantly improves evaluation reliability compared to traditional evaluations with scalar scoring.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!