코딩 에이전트를 위한 베이지안 제어
Bayesian control for coding agents
최신 코딩 에이전트는 LLM 생성기와 저렴한 진단 도구 및 고가의 검증 도구를 결합합니다. 이러한 도구 사용 결정은 일반적으로 정해진 규칙을 사용하는 오케스트레이터에 의해 관리되지만, 불확실성을 고려하지 않는 경우가 많습니다. 본 연구에서는 오케스트레이션을 비용 민감적인 순차적 가설 검정으로 정의하며, 베이지안 제어기는 후보 코드의 정확성에 대한 믿음을 유지하고, 추가 증거 수집 여부, 후보 코드 개선 여부, 검증 여부 또는 중단 여부를 동적으로 결정합니다. 6개의 생성 모델과 9개의 코딩 벤치마크를 통해 실험한 결과, 베이지안 제어는 검증 비용이 높고 비평가가 유용하지만 완벽하지 않은 경우에 가장 큰 효과를 발휘하는 것으로 나타났습니다. 또한, 제어 기능 외에도, 믿음 상태는 해석 가능한 정확도 점수를 제공하며, 이는 불확실성 측정 측면에서 토큰 확률 및 도구 성공률을 기준으로 한 기존 방법보다 우수한 성능을 보입니다.
Modern coding agents pair LLM generators with various tools, including cheap diagnostics and expensive verifiers. The tool-use decisions are typically governed by orchestrators that often use fixed rules and ignore uncertainty. We formulate orchestration as cost-sensitive sequential hypothesis testing: a Bayesian controller maintains a belief over candidate correctness and dynamically decides whether to gather more evidence, refine the candidate, verify it, or stop. Across six generators and nine coding benchmarks, Bayesian control proves to be most valuable when verification is costly and critics are informative but imperfect. Beyond control, the belief state yields an interpretable correctness score that outperforms token-probability and raw tool-success baselines for uncertainty quantification.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.