2607.11346v1 Jul 13, 2026 cs.AI

컴파일 후 페이지 처리: 실행 가능한 SOP 프로그램 및 절차형 LLM 에이전트를 위한 기능 기반 런타임

Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents

Chenglin Yu
Chenglin Yu
Citations: 36
h-index: 4
Ming Li
Ming Li
Citations: 15
h-index: 2
Lichao Yin
Lichao Yin
Citations: 20
h-index: 1
Ying Yu
Ying Yu
Citations: 28
h-index: 3
Qingxin Fan
Qingxin Fan
Citations: 0
h-index: 0
RunyangRay Zhong
RunyangRay Zhong
Citations: 0
h-index: 0

기업용 에이전트는 장기적이고 조건부이며 안전에 중요한 표준 운영 절차(SOP)를 준수해야 합니다. 본 연구에서는 기계가 읽을 수 있는 SOP 제약을 실행 가능한 유사 코드로 컴파일하고, LLM이 의미론적 실행을 수행하는 동안 프로그램 가이드(PG) 스택 머신을 사용하여 활성 프레임을 페이지 처리합니다. 6개의 모델에 대한 세 가지 그룹의 SOPBench 연구를 통해 표현과 런타임 간의 차이를 분석한 결과, 컴파일된 텍스트가 성능 저하를 초래하는 경우는 거의 없으며, 공식적인 문장 표현이 상대적으로 부족한 경우 최대 16.0점의 성능 향상을 보였습니다. 런타임 가이드 기능은 기능 기반으로 제한됩니다. 두 개의 강력한 모델에서 독립적으로 긍정적인 PG 효과가 나타나는 것을 확인했으며(58:19 및 75:31의 불일치 쌍), 반면 약한 모델에서는 성능 저하가 발생했습니다. 전체 프로그램 커서 제거 실험(활성 프레임 먼저, 전체 프로그램은 유지)을 통해 강력한 모델에서 얻는 거부율 향상의 상당 부분을 회복할 수 있으며, 선택적 가시성은 추가적인 개선 효과를 제공합니다. 페어링된 프로브 및 감사 측정값을 통해 이러한 차이가 재구성 능력보다는 자발적인 상태 제어를 반영한다는 것을 확인했습니다. 'Bank' 데이터셋에 대한 실험 결과, 세 가지 그룹의 성능이 70.4에서 86.4로, 그리고 92.8로 향상되었으며, 거부율 정확도는 100%를 달성했습니다. 실용적인 가이드라인은 다음과 같습니다: 먼저 컴파일을 수행하고, 모델 수준의 제어 기능 확인 후 활성 프레임 페이지 처리를 활성화합니다.

Original Abstract

Enterprise agents must follow long-horizon, conditional, safety-critical standard operating procedures (SOPs). We compile machine-readable SOP constraints into executable pseudo-code and run them with a program-guided (PG) stack machine that pages the active frame while an LLM performs semantic execution. A three-arm SOPBench study across six models separates representation from runtime: compiled text never significantly hurts and gains up to 16.0 points where official prose underperforms. Runtime guidance is capability-gated. Two strong models independently show positive seven-domain PG contrasts (58:19 and 75:31 discordant pairs), whereas weak models are harmed. A full-program cursor ablation (active frame first, complete program retained) recovers much of the strong-model refusal gain; selective visibility adds a smaller improvement. Paired probe and audit measurements track this divide to spontaneous state discipline rather than reconstruction ability. On Bank the three primary arms rise from 70.4 to 86.4 to 92.8, with 100% refusal correctness. Practical guidance: compile first; enable active-frame paging only after a model-level discipline check.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!