2607.28587v1 Jul 30, 2026 cs.SE

PAIChecker: SWE-Bench과 유사한 벤치마크에서 PR(Pull Request)과 이슈 불일치 현상 분석 및 검증

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks

Junjielong Xu
Junjielong Xu
The Chinese University of Hong Kong, Shenzhen
Citations: 381
h-index: 8
Pinjia He
Pinjia He
Citations: 63
h-index: 3
Manyi Wang
Manyi Wang
Citations: 29
h-index: 3

SWE-bench와 유사한 벤치마크는 LLM의 문제 해결 능력을 평가하는 데 널리 사용됩니다. 이들은 일반적으로 다음과 같은 공통적인 구성 과정을 따릅니다: 각 PR은 PR 설명에서 추출된 이슈 참조를 통해 연결된 이슈와 쌍을 이루며, 이슈 설명이 문제 정의로, PR 패치는 테스트 오라클로 사용됩니다. 그러나 대규모 저장소를 개발하고 유지 관리하는 데 내재된 복잡성으로 인해 실제로는 이러한 PR-이슈 페어링이 종종 불일치되는 경우가 많습니다. 본 연구에서는 SWE-bench Verified 인스턴스를 체계적으로 분석하여 13.6%가 11가지 세분화된 시나리오에서 5가지 패턴을 통해 불일치를 보이는 것을 확인했습니다. 향후 이러한 벤치마크의 신뢰성과 확장성을 확보하기 위해, 우리는 SWE-bench와 유사한 벤치마크에서 PR-이슈 불일치 현상을 검증하는 다중 에이전트 시스템인 PAIChecker를 제안합니다. 특히, PAIChecker는 특정 패턴 식별, 교차 에이전트 라벨 합성, 코드 수준 검증을 결합한 3단계 설계를 채택하여 더욱 정확하고 일반화 가능하며 점진적으로 검증된 탐지를 가능하게 합니다. SWE-Gym 및 SWE-bench Multilingual에 대한 실험 결과, PAIChecker는 네 가지 LLM 백본 모두에서 최상의 성능을 달성했으며, 각각 92.12%와 91.67%의 이진 정확도를 기록했습니다.

Original Abstract

SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability. They typically follow a common construction pipeline: each PR (Pull Request) is paired with its linked issue by extracting issue references from the PR description; the issue description is used as the problem statement, and the PR patch serves as the test oracle. However, due to the inherent complexity of developing and maintaining large repositories, such PR-Issue pairings are often misaligned in practice. In this work, we systematically study SWE-bench Verified instances, finding that 13.6% exhibit misalignment across five patterns in eleven fine-grained scenarios. To enable reliable and scalable construction of those benchmarks in the future, we propose PAIChecker, a multi-agent system for checking PR-Issue misalignment in SWE-bench-like benchmarks. Specifically, PAIChecker adopts a three-phase design that combines specific pattern identification, cross-agent label synthesis, and code-level validation, thereby enabling more accurate, generalizable, and progressively verified detection. Experiments on SWE-Gym and SWE-bench Multilingual show that PAIchecker achieves the best performance across all four LLM backbones, reaching up to 92.12% and 91.67% binary accuracy, respectively.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!