2601.21064v2 Jan 28, 2026 cs.LG

심층 복합 AI 시스템을 위한 텍스트 균형 전파

Textual Equilibrium Propagation for Deep Compound AI Systems

Minghui Chen
Minghui Chen
Citations: 70
h-index: 6
Wenlong Deng
Wenlong Deng
Citations: 205
h-index: 7
James Zou
James Zou
Citations: 5
h-index: 1
Han Yu
Han Yu
Citations: 3
h-index: 1
Xiaoxiao Li
Xiaoxiao Li
Citations: 154
h-index: 5

대규모 언어 모델(LLM)은 점점 더 복합 AI 시스템의 일부로 사용되는데, 이러한 시스템은 여러 모듈(예: 검색기, 도구, 검증기)을 긴 작업 흐름에 걸쳐 조정합니다. 최근에는 텍스트 피드백을 전역적으로 전파하는 방식(예: TextGrad)이 등장하여 이러한 파이프라인을 최적화하는 데 도움이 되지만, 시스템의 깊이가 깊어짐에 따라 성능이 저하되는 것을 확인했습니다. 특히, 장기적인 에이전트 기반 작업 흐름은 두 가지 깊이 확장 관련 문제점을 보입니다. 첫째, 텍스트 그래디언트가 깊이에 따라 기하급수적으로 증가하여 메시지 길이가 지나치게 길어지고, 평가 편향을 증폭시킵니다. 둘째, 제한된 장기 컨텍스트 처리 능력으로 인해 모델이 부분적인 피드백에 과도하게 집중하고, 긴 피드백을 압축하는 과정에서 하위 메시지의 구체성이 점진적으로 감소하여 상위 단계로 전달됩니다. 이러한 문제점을 해결하기 위해, 에너지 기반 모델의 균형 전파(Equilibrium Propagation)에서 영감을 받은 로컬 학습 원리인 텍스트 균형 전파(Textual Equilibrium Propagation, TEP)를 제안합니다. TEP는 두 단계로 구성됩니다. 첫째, 로컬 LLM 평가기가 균형 상태에 도달할 때까지(더 이상 개선이 없을 때) 프롬프트를 반복적으로 개선하는 '자유 단계'입니다. 둘째, '강제 단계'로, 작업 수준의 목표를 사용하여 순방향 신호 방식으로 전파하면서, 제한적인 수정 강도를 가진 근접 프롬프트 편집을 적용합니다. 이러한 설계는 로컬 프롬프트 최적화를 가능하게 하고, 전역적인 텍스트 역전파의 계산 부담과 신호 저하 없이, 제어된 방식으로 전역 목표에 적응하도록 지원합니다. 장기 QA 벤치마크 및 다중 에이전트 도구 사용 데이터셋에서 TEP는 TextGrad와 같은 전역 전파 방식보다 일관되게 정확도와 효율성을 향상시킵니다. 이러한 성능 향상은 시스템 깊이가 깊어질수록 더욱 두드러지며, 심층 복합 AI 시스템에서 블랙박스 LLM 구성 요소의 실용성을 유지합니다.

Original Abstract

Large language models (LLMs) are increasingly deployed as part of compound AI systems that coordinate multiple modules (e.g., retrievers, tools, verifiers) over long-horizon workflows. Recent approaches that propagate textual feedback globally (e.g., TextGrad) make it feasible to optimize such pipelines, but we find that performance degrades as system depth grows. In particular, long-horizon agentic workflows exhibit two depth-scaling failure modes: 1) exploding textual gradient, where textual feedback grows exponentially with depth, leading to prohibitively long message and amplifies evaluation biases; and 2) vanishing textual gradient, where limited long-context ability causes models overemphasize partial feedback and compression of lengthy feedback causes downstream messages to lose specificity gradually as they propagate many hops upstream. To mitigate these issues, we introduce Textual Equilibrium Propagation (TEP), a local learning principle inspired by Equilibrium Propagation in energy-based models. TEP includes two phases: 1) a free phase where a local LLM critics iteratively refine prompts until reaching equilibrium (no further improvements are suggested); and 2) a nudged phase which applies proximal prompt edits with bounded modification intensity, using task-level objectives that propagate via forward signaling rather than backward feedback chains. This design supports local prompt optimization followed by controlled adaptation toward global goals without the computational burden and signal degradation of global textual backpropagation. Across long-horizon QA benchmarks and multi-agent tool-use dataset, TEP consistently improves accuracy and efficiency over global propagation methods such as TextGrad. The gains grows with depth, while preserving the practicality of black-box LLM components in deep compound AI system.

1 Citations
0 Influential
3.5 Altmetric
18.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!