능력 기반 계획: 목표 달성을 위한 비용 발견 및 근시안적 실험 선택의 한계
Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment Selection
과학적 발견을 자동화하는 시스템은 반복적으로 어떤 실험을 수행할지, 어떤 가설을 검증할지, 어떤 도구를 개발할지, 그리고 언제 중단할지를 결정해야 합니다. 많은 시스템이 기대 정보 획득량 대비 비용 또는 학습된 가능성 점수와 같은 근시안적 지표를 최대화하여 이러한 결정을 내립니다. 우리는 이 접근 방식의 구조적인 한계를 밝혀냈습니다. 어떤 행동들은 건설적입니다. 즉, 즉시 얻는 정보가 아닌 미래에 수행할 수 있는 행동을 가능하게 하는 인식적 능력을 (기구, 분석법, 파이프라인, 시뮬레이터 또는 추상화) 획득합니다. 확신을 갖게 될 때까지의 최저 비용 경로가 이러한 일련의 건설 단계를 필요로 할 때, 특정 범위 내에서만 정보를 얻을 수 있는 행동만을 평가하는 계획기는 첫 번째 건설 단계를 제대로 평가할 수 없습니다. 이는 해당 범위 내에서 아무런 정보도 제공하지 않으며, 작은 양의 정보라도 제공하는 측정에 의해 항상 능가되기 때문입니다. 우리는 목표 지향적 발견을 신념 공간에서의 확률적 최단 경로 문제로 공식화하며, 건설적인 실험은 이후 행동 그래프를 변경한다는 점을 고려합니다. 또한 모든 탐색 깊이 d에 대해, 모든 근시안적 정보 최대화 계획기가 무한한 근사 비율을 갖는 경우와 목표 지점에 도달하지 못하는 관련 경우 모두 존재함을 증명했습니다. 이 메커니즘은 '능력 구별 불가능성' 정리에 의해 작동합니다. 즉, 특정 범위 내에서 능력을 획득하는 것은 관찰적으로 아무런 행동도 수행하지 않는 것과 구별할 수 없습니다. 이는 능력 기반 계획이 곡률(하위 모듈성) 및 정보 순서(적응성 격차)와는 다른 '접근 가능성'의 어려움이라는 점을 시사합니다. 우리는 능력을 고려하는 비용-목표 휴리스틱 h = h_cap + h_exp를 사용하는 점진적인 재계획기인 CG-Plan을 소개했습니다. 통제된 테스트 환경에서 성능 격차는 능력 기반 제한 조건 하에서만 나타나며, 모든 고정된 범위에 대해 지속되며, 데이터 일관성을 갖는 제안자로부터 '거의 맞춤' 가설이 나올 때 발생합니다.
Systems that automate scientific discovery must repeatedly decide which experiment to run, which hypothesis to test, which tool to build, and when to stop. Many systems make these decisions by maximizing a myopic score such as expected information gain per unit cost or a learned plausibility score. We identify a structural limitation of this approach. Some actions are constructive: they acquire an epistemic capability (an instrument, assay, pipeline, simulator, or abstraction) whose value lies not in the information returned immediately but in the future actions it makes available. When the least-cost route to a confident answer requires a chain of such constructions, a planner that scores actions only by information obtainable within a bounded horizon cannot value the first construction: it yields no information within the horizon and is dominated by any measurement with positive information, however small. We formulate goal-directed discovery as a stochastic shortest-path problem in belief space in which constructive experiments change the downstream action graph, and prove that for every lookahead depth d there is an instance on which every myopic information-maximizing planner has an unbounded approximation ratio, and a related instance on which it never reaches the goal. The mechanism is a capability-indistinguishability lemma: within the horizon, acquiring a capability can be observationally indistinguishable from paying for a null action. This establishes capability gating as a reachability axis of difficulty distinct from curvature (submodularity) and information order (adaptivity gaps). We introduce CG-Plan, an incremental replanner with a capability-aware cost-to-go heuristic h = h_cap + h_exp. In a controlled testbed, the performance gap appears only under gating, persists for every fixed horizon, and arises when near-miss hypotheses come from a data-consistent proposer.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.