2606.24311v1 Jun 23, 2026 cs.AI

LemonHarness 기술 보고서

LemonHarness Technical Report

K. Ren
K. Ren
Citations: 1
h-index: 1
Zimo Yin
Zimo Yin
Citations: 3
h-index: 1
Conglin Yin
Conglin Yin
Citations: 12
h-index: 2
Zeping Chen
Zeping Chen
Citations: 1
h-index: 1
Yanzhi Xu
Yanzhi Xu
Citations: 830
h-index: 14
Jianpin Fan
Jianpin Fan
Citations: 129
h-index: 7
Jiawei Liu
Jiawei Liu
Citations: 40
h-index: 3
Ming He
Ming He
Citations: 108
h-index: 5
Lei Zhang
Lei Zhang
Citations: 61
h-index: 4
Liu Yang
Liu Yang
Citations: 48
h-index: 2
Yixuan Wu
Yixuan Wu
Citations: 2
h-index: 1
Yunchen Huo
Yunchen Huo
Citations: 2
h-index: 1
Fubo Sun
Fubo Sun
Citations: 1
h-index: 1
Jiachen Liu
Jiachen Liu
Citations: 0
h-index: 0
Jiaying Li
Jiaying Li
Citations: 0
h-index: 0
Yubin Huangfu
Yubin Huangfu
Citations: 99
h-index: 4
Ronghua Li
Ronghua Li
Citations: 0
h-index: 0
X. Su
X. Su
Citations: 2,659
h-index: 25
Likang Wu
Likang Wu
Citations: 9
h-index: 2
Hongke Zhao
Hongke Zhao
Citations: 10
h-index: 1
Xiaohui Geng
Xiaohui Geng
Citations: 28
h-index: 2

대규모 언어 모델(LLM) 에이전트가 더 복잡하고 긴 작업을 수행함에 따라, 작업 과정에서 여러 단계에 걸쳐 작업 공간의 상태가 점차적으로 변경됩니다. 그러나 일반적으로 에이전트는 도구의 출력과 로그 조각만을 관찰하며, 실제 상태 변화는 파일 시스템에서 발생합니다. 명시적인 작업 공간 경계가 없으면 파일 쓰기와 같은 상태를 변경하는 작업이나 임시 결과물 생성 등으로 인해 변경 사항이 여러 경로에 흩어질 수 있습니다. 시간이 지남에 따라 이러한 제약이 느슨한 변경 사항들이 누적되면서 수정된 파일과 같이 상태를 추적하기 어려워집니다. 본 논문에서는 장기적인 작업을 위한 통합 실행 프레임워크인 LemonHarness를 소개합니다. LemonHarness는 명확하게 정의된 작업 공간 내에서 상태 변경 작업을 제한함으로써 명시적인 실행 경계를 설정하고, 모델 호출, 도구 실행 및 규칙 기반 지식을 단일 제어된 경계 내에 포함시킵니다. 파일 쓰기, 의존성 설치 및 임시 결과물 생성과 같은 상태를 변경하는 작업은 구조화된 도구 인터페이스를 통해 수행되며, 실행 피드백은 이후 모델의 의사 결정에 사용될 수 있는 관찰 데이터로 기록됩니다. 또한, LemonHarness는 재사용 가능한 규칙 기반 지식 베이스를 도입하여 반복적인 실행 규칙과 승인 기준을 런타임 지식으로 변환합니다. 더 나아가, LemonHarness는 시간 인식을 갖춘 실행 메커니즘을 추가하여 모델에게 경과 시간 및 남은 시간을 알려줌으로써, 시간이 촉박해지거나 과도한 검증이 필요한 경우 탐색, 구현 및 검증 노력을 재조정하고 타임아웃을 방지할 수 있도록 합니다. Terminal-Bench 2.0에서 LemonHarness_GPT-5.3-CodeX는 445번의 시도에서 84.49%의 정확도를 달성했으며, 동일한 프레임워크를 더 강력한 GPT-5.5 모델과 결합하면 평균 정확도가 86.52%로 향상되었습니다. 이러한 결과는 통합된 런타임 경계, 호출 가능한 규칙 기반 지식 및 시간 인지 실행이 장기적인 에이전트 실행의 안정성을 향상시킬 수 있음을 시사합니다.

Original Abstract

As large language model (LLM) agents are applied to longer tasks, they increasingly modify workspace state across multiple rounds of iteration. However, agents typically observe only tool outputs and log fragments, while the actual state changes occur in the file system. Without explicit workspace boundaries, state-changing operations such as file writes and temporary artifact generation may scatter changes across paths. Over time, these weakly constrained changes accumulate, making states such as modified files difficult to track. This paper presents LemonHarness, an integrated execution framework for long-horizon agents. LemonHarness establishes an explicit execution boundary by constraining state-changing operations within a clearly defined workspace and bringing model invocation, tool execution, and rule knowledge within a single controlled boundary. State-changing operations, including file writes, dependency installation, and temporary artifact creation, are executed through structured tool interfaces, with execution feedback recorded as observations available to subsequent model decisions. The system also introduces a reusable rule knowledge base, which turns recurring execution rules and acceptance criteria into runtime knowledge. LemonHarness further adds a time-aware execution mechanism that exposes elapsed and remaining budget to the model, so it can rebalance exploration, implementation, and validation effort as time pressure shifts and avoid timeouts from long waits or excessive verification. On Terminal-Bench 2.0, LemonHarness_GPT-5.3-CodeX reached 84.49% accuracy over 445 trials; pairing the same framework with the stronger GPT-5.5 backbone raised the average accuracy to 86.52% across five jobs. The results suggest that a unified runtime boundary, callable rule knowledge, and time-aware execution can improve the stability of long-horizon agent execution.

0 Citations
0 Influential
12.5 Altmetric
62.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!