제어 효과: 오케스트레이션 설계가 엔터프라이즈 에이전트형 AI의 토큰 경제를 어떻게 결정하는가
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
오늘날 에이전트형 AI 개발은 '토큰 최대화' 방식으로 이루어집니다. 즉, 토큰을 사용하여 추론 과정의 길이, 턴 수, 도구 활용 범위, 재활용 컨텍스트 크기를 늘리므로, 작업당 필요한 토큰 수는 작업 가치보다 빠르게 증가합니다. 토큰 가격 하락은 이러한 현상을 가리고 있지만, 총 지출은 계속 증가하고 있습니다. 우리는 '제어(harness)'가 토큰 최대화에 대한 결정적인 해결책이라고 주장합니다. 제어는 컨텍스트를 구성하고, 도구를 노출하며, 턴 순서를 관리하고, 작업을 위임하며, 엔터프라이즈 수준의 관찰 및 거버넌스를 제공하는 오케스트레이션 계층입니다. 우리는 통제된 실험을 통해 이를 입증했습니다. 구체적으로, 22개의 평가 작업과 6가지 기본 모델(Claude Sonnet 4.6, Gemini 3.1, Gemini Flash 3.5, Qwen 3.6, GLM 5.1, Palmyra X6)을 사용하여 오케스트레이션 계층만 변경했습니다. 기존의 일반적인 생산 루프와 'Writer Agent Harness'를 비교한 결과, 모델을 동일하게 유지하면서 제어가 작업당 평균 비용을 41% (0.21달러에서 0.12달러로), 중앙 처리 시간을 44% (48초에서 27초로), 작업당 토큰 수를 38% (14,200개에서 8,800개로) 감소시켰으며, 작업 완료 품질은 동일하게 유지되었습니다 (0.78에서 0.81로). 효율성은 모델에 관계없이 모든 모델에서 비용을 절감했습니다 (33-61%). 반면, 품질 향상은 모델의 성능에 따라 달라졌습니다. 모델의 성능이 좋을수록 제어를 통해 얻는 품질 향상 효과가 더 컸으며, 그 상관관계는 거의 완벽했습니다 (r=0.99, n=6). 이를 우리는 '제어 레버리지'라고 명명합니다. 달러당 작업 품질은 82% 증가했으며, 백만 개의 토큰당 완료된 작업 수는 54.9개에서 92.0개로 증가했습니다. 이 워크로드에서는 오케스트레이션 계층이 모델 선택보다 작업당 비용에 더 큰 영향을 미쳤습니다. 우리는 오케스트레이션 계층에서의 토큰 경제를 공식화하고 (프롬프트 캐싱을 통한 효과적인 입력 가격 포함), 제어 효과의 기반이 되는 6가지 메커니즘 그룹 (캐시 관리부터 오류 처리까지)을 자세히 설명하고, 널리 사용되는 6가지 에이전트 시스템을 동일한 기준으로 비교합니다. 또한, 오케스트레이션 계층의 효율성은 조직에서 사용하는 모든 모델에 걸쳐 증폭되는 유일한 요소이며, 현재 및 미래 모두에 해당한다고 주장합니다.
Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token prices mask the pattern; total spend rises anyway. We argue the decisive lever against token maxing is the harness: the orchestration layer that assembles context, exposes tools, sequences turns, delegates work, and carries enterprise observability and governance. We isolate it with a controlled swap: 22 locked evaluation tasks, six foundation models (Claude Sonnet 4.6, Gemini 3.1, Gemini Flash 3.5, Qwen 3.6, GLM 5.1, Palmyra X6), changing only the orchestration layer -- a frozen conventional production loop versus the Writer Agent Harness. Holding models constant, the harness cuts blended cost per task 41% ($0.21->$0.12), median wall-clock 44% (48s->27s), and tokens per task 38% (14.2k->8.8k), with task-completion quality at parity (0.78->0.81, directional at this sample size). Efficiency is model-invariant -- every model gets cheaper (33-61%) -- while quality gains are capability-dependent: a model's gain correlates almost perfectly with its baseline strength (r=0.99, n=6), a phenomenon we term harness leverage. Quality per dollar rises 82%; task-completions per million tokens rise from 54.9 to 92.0. On this workload the orchestration layer moved cost per task more than the full spread of the model menu did. We formalize token economics at the orchestration layer (including effective input price under prompt caching), detail the six mechanism families behind the effect -- cache-shape discipline to failure-spend governance -- compare six widely used agent systems on the same axes, and argue the harness is the one component whose efficiency multiplies across every model an organization runs -- present and future.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.