Σ-Mem: LLM 기반 다중 에이전트 시스템을 위한 온라인 신뢰성 메모리
$Σ$-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
메모리는 장기적인 작업을 수행하는 LLM (대규모 언어 모델) 에이전트에 있어 핵심적인 역할을 합니다. 그러나 기존의 메모리 시스템은 주로 상호 작용 내용을 보존하는 데 중점을 두며, 어떤 에이전트를 신뢰할 수 있고 어떤 조건에서 신뢰할 수 있는지 모델링하지 못합니다. 이러한 제한 사항은 특히 다중 에이전트 시스템에서 중요한 문제이며, 중앙 모델이 직접적으로 동료 에이전트의 응답의 타당성 또는 상관관계를 검증하기 어려울 수 있습니다. 본 논문에서는 개별 에이전트의 과거 능력 증거와 전체 에이전트 그룹 간의 관계 증거를 기록하는 온라인 신뢰성 메모리인 $Σ$-Mem을 소개합니다. 이러한 모든 증거는 실수형 대칭 상태로 유지되며, 의사 결정 후 제공되는 정확도 피드백을 통해 업데이트됩니다. Weyl 부등식에 따르면, 각 이벤트 수준에서의 업데이트로 인해 발생하는 스펙트럼 변화는 제한적이며, 이를 통해 기본 모델을 재학습하지 않고도 안정적인 온라인 적응이 가능합니다. $Σ$-Mem은 일반적인 쓰기 및 읽기 인터페이스를 제공하며, 동일한 메모리는 중앙 모델의 추가 제어, 응답과 독립적인 동료 에이전트 라우팅 또는 신뢰성 가중 투표에 사용될 수 있습니다. Qwen 패밀리의 5가지 모델에서 $Σ$-Mem은 반사실적 신뢰성 변화에 적응하고, 새로운 에이전트 및 작업 영역으로 일반화됩니다. 또한 메모리 직접 읽기는 다수 투표 방식과 최상의 고정된 동료 에이전트를 사용하는 경우보다 전체 OOD (Out-of-Distribution) 평가 세트에서 더 우수한 성능을 보입니다. 더욱이, 정확도 피드백이 많아질수록 성능이 꾸준히 향상되는 것으로 나타나, $Σ$-Mem이 점진적으로 실행 가능한 신뢰성 정보를 축적한다는 것을 시사합니다. 이러한 결과는 신뢰성 메모리를 LLM 기반 다중 에이전트 시스템에서 적응적인 조화를 위한 재사용 가능한 기반으로 확립합니다.
Memory is central to long-horizon LLM agents, yet existing memory systems primarily preserve interaction content rather than modeling which agents can be trusted and under what conditions. This limitation is particularly important in multi-agent systems, where a central model may be unable to directly verify plausible or correlated peer responses. We introduce $Σ$-Mem, an online reliability memory that records historical competence evidence for individual peers and peer relationship evidence across the peer set. Both forms of evidence are maintained as real symmetric states and updated from post-decision correctness feedback. By Weyl's inequality, the spectral change caused by each event-level update is bounded, enabling stable online adaptation without retraining the underlying models. $Σ$-Mem provides a general write-and-read interface: the same memory can be used for residual steering of a central model, response-free peer routing, or reliability-weighted voting. Across five Qwen-family models, $Σ$-Mem adapts to counterfactual reliability shifts and generalizes to unseen peers and task domains. Direct memory readouts also outperform majority voting and the best fixed peer over the full OOD evaluation set. Moreover, performance improves consistently as more correctness feedback becomes available, indicating that $Σ$-Mem progressively accumulates actionable reliability information. These results establish reliability memory as a reusable foundation for adaptive coordination in LLM-based multi-agent systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.