2602.02262v2 Feb 02, 2026 cs.SE

OmniCode: 소프트웨어 엔지니어링 에이전트 평가를 위한 벤치마크

OmniCode: A Benchmark for Evaluating Software Engineering Agents

Claas Beger
Claas Beger
Citations: 23
h-index: 2
Gloria Geng
Gloria Geng
Citations: 2
h-index: 1
Atharv Sonwane
Atharv Sonwane
Citations: 390
h-index: 7
Eng-Shen Tu
Eng-Shen Tu
Citations: 20
h-index: 2
Wei-Chung Lu
Wei-Chung Lu
Citations: 3
h-index: 1
Carter Larsen
Carter Larsen
Citations: 46
h-index: 2
Debjit Dhar
Debjit Dhar
Citations: 2
h-index: 1
Rachel Chen
Rachel Chen
Citations: 2
h-index: 1
Ronit Pattanayak
Ronit Pattanayak
Citations: 2
h-index: 1
T. Dang
T. Dang
Citations: 5
h-index: 2
Guohao Chen
Guohao Chen
Citations: 2
h-index: 1
Kevin Ellis
Kevin Ellis
Citations: 7
h-index: 2
Saikat Dutta
Saikat Dutta
Citations: 14
h-index: 2

LLM 기반 코딩 에이전트는 실제 소프트웨어 개발 방식을 혁신하고 있습니다. 더 나은 코딩 에이전트 개발을 촉진하기 위해서는 다양한 소프트웨어 엔지니어링 작업을 엄격하게 평가할 수 있는 도전적인 벤치마크가 필요합니다. 그러나 HumanEval 및 SWE-Bench와 같은 기존 코딩 벤치마크는 주로 경쟁 프로그래밍 및 패치 생성과 같이 제한적인 작업에 초점을 맞추고 있습니다. 실제로는 소프트웨어 엔지니어는 실제 소프트웨어 개발에 필요한 더 광범위한 작업을 처리해야 합니다. 이러한 격차를 해소하기 위해, 코드 또는 패치 생성 외에 더 넓고 다양한 작업 범주를 포함하는 새로운 소프트웨어 엔지니어링 벤치마크인 OmniCode를 제안합니다. OmniCode는 총 1794개의 작업으로 구성되며, Python, Java, C++ 세 가지 프로그래밍 언어와 버그 수정, 테스트 생성, 코드 리뷰 수정, 스타일 수정의 네 가지 주요 범주를 포함합니다. 기존 소프트웨어 엔지니어링 벤치마크와 달리, OmniCode의 작업은 (1) 모호한 문제를 제거하기 위해 수동으로 검증되었으며, (2) 데이터 유출 문제를 방지하기 위해 인공적으로 생성되었거나 최근에 큐레이션되었습니다. 이를 통해 제한된 실제 데이터를 기반으로 다양한 소프트웨어 작업을 인공적으로 생성하는 새로운 프레임워크를 제공합니다. OmniCode를 SWE-Agent와 같은 인기 있는 에이전트 프레임워크로 평가한 결과, Python의 버그 수정 작업에서는 좋은 성능을 보이지만, 테스트 생성 작업 및 C++ 및 Java와 같은 언어에서는 성능이 부족하다는 것을 확인했습니다. 예를 들어, SWE-Agent는 DeepSeek-V3.1을 사용하여 Java 테스트 생성 작업에서 최대 20.9%의 정확도를 달성했습니다. OmniCode는 견고한 벤치마크 역할을 하고 소프트웨어 개발의 다양한 측면에서 뛰어난 성능을 보이는 에이전트 개발을 촉진하는 것을 목표로 합니다. 코드 및 데이터는 https://github.com/seal-research/OmniCode 에서 확인할 수 있습니다.

Original Abstract

LLM-powered coding agents are redefining how real-world software is developed. To drive the research towards better coding agents, we require challenging benchmarks that can rigorously evaluate the ability of such agents to perform various software engineering tasks. However, popular coding benchmarks such as HumanEval and SWE-Bench focus on narrowly scoped tasks such as competition programming and patch generation. In reality, software engineers have to handle a broader set of tasks for real-world software development. To address this gap, we propose OmniCode, a novel software engineering benchmark that contains a broader and more diverse set of task categories beyond code or patch generation. Overall, OmniCode contains 1794 tasks spanning three programming languages (Python, Java, and C++) and four key categories: bug fixing, test generation, code review fixing, and style fixing. In contrast to prior software engineering benchmarks, the tasks in OmniCode are (1) manually validated to eliminate ill-defined problems, and (2) synthetically crafted or recently curated to avoid data leakage issues, presenting a new framework for synthetically generating diverse software tasks from limited real-world data. We evaluate OmniCode with popular agent frameworks such as SWE-Agent and show that while they may perform well on bug fixing for Python, they fall short on tasks such as Test Generation and in languages such as C++ and Java. For instance, SWE-Agent achieves a maximum of 20.9% with DeepSeek-V3.1 on Java Test Generation tasks. OmniCode aims to serve as a robust benchmark and spur the development of agents that can perform well across different aspects of software development. Code and data are available at https://github.com/seal-research/OmniCode.

3 Citations
0 Influential
35.92453324894 Altmetric
13.9 Score
Original PDF
11

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!