2605.08013v1 May 08, 2026 cs.AI

선택적 관찰 환경에서의 구조화된 행동 기반 보상 부여를 통한 CLI 에이전트 학습

Learning CLI Agents with Structured Action Credit under Selective Observation

Haoyang Su
Haoyang Su
Citations: 17
h-index: 3
Ying Wen
Ying Wen
Citations: 18
h-index: 2

명령줄 인터페이스(CLI) 에이전트는 진화하는 파일 시스템, 실행 가능한 명령줄 프로그램 및 온라인 실행 피드백을 통한 에이전트-컴퓨터 상호 작용을 위한 실용적인 패러다임으로 부상하고 있습니다. 최근 연구에서는 강화 학습(RL)을 사용하여 검증 가능한 작업 피드백으로부터 이러한 상호 작용 능력을 학습하는 방법을 사용했지만, CLI 액션의 고유한 구조화된 속성을 학습 신호로 활용하는 방법은 드뭅니다. 이러한 활용되지 않은 액션 구조 외에도, CLI 학습은 에이전트 코딩에 대한 두 가지 주요 병목 현상을 야기합니다. 첫째, 에이전트는 부분적인 관찰을 통해 대규모 코드베이스에서 작업과 관련된 증거를 식별해야 합니다. 둘째, 희소한 최종 보상은 긴 다단계 경로를 형성하는 행동에 할당되어야 합니다. 우리는 셸 기반 정보 추출 및 파일 편집 작업을 통해 이러한 병목 현상을 연구합니다. 선택적 관찰을 위해, 동일한 CLI에 대한 토큰 예산을 고려하여 컨텍스트를 선택하는 추론 시간 메커니즘인 σ-Reveal을 도입합니다. 보상 할당을 위해, 표준 에이전트 기반 RL의 알고리즘 복잡성을 유지하는 에이전트 기반 RL 방법인 Action Advantage Assignment (A³ )을 제안합니다. A³는 에피소드 수준의 상대적 피드백, 추상 구문 트리(AST) 기반 액션 서브 체인 잔차 및 트리 수준의 경로 마진으로부터 턴 수준의 이점을 구성합니다. 또한, 이 문제 설정을 더욱 평가하기 위해, 저장소 환경에서 CLI 작업을 다루는 검증 가능한 데이터셋인 ShellOps를 구축했습니다.

Original Abstract

Command line interface (CLI) agents are emerging as a practical paradigm for agent-computer interaction over evolving filesystems, executable command line programs, and online execution feedback. Recent work has used reinforcement learning (RL) to learn these interaction abilities from verifiable task feedback, yet few methods exploit the native structured attributes of CLI actions as learning signals. Beyond this underused action structure, CLI learning also couples two bottlenecks for coding agents. First, the agent must identify task-relevant evidence in a large codebase from partial observations. Second, sparse terminal rewards must be assigned to the actions that shape a long multi-turn trajectory. We study these bottlenecks through shell-driven information extraction and file editing tasks. For selective observation, we introduce $σ$-Reveal, an inference-time mechanism that selects token-budgeted context for the same CLI. For credit assignment, we propose Action Advantage Assignment ($\mathrm{A}^3$), a native agentic RL method that preserves the algorithmic complexity of standard agentic RL. $\mathrm{A}^3$ constructs turn-level advantages from episode-level relative feedback, abstract syntax tree (AST) based action sub-chain residuals, and tree-level trajectory margins. To further evaluate this problem setting, we construct ShellOps, a verifiable dataset suite covering CLI tasks in repository environments.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!