기억해야 할 것을 학습하기: 제약 조건 최적화를 통한 장기 언어 에이전트를 위한 관찰 가능성을 고려한 안전한 메모리 유지
Learning What to Remember: Observability-Safe Memory Retention via Constrained Optimization for Long-Horizon Language Agents
장기적인 언어 에이전트는 관찰 결과, 추론 과정 및 검색된 사실들을 축적하며, 이는 종종 제한된 컨텍스트 창을 초과하게 됩니다. 따라서 메모리 유지는 기본적인 자원 할당 문제입니다. 기존의 메모리 시스템은 휴리스틱 기반 점수 부여, 검색 최적화 또는 학습 기반 압축을 통해 관리를 개선하지만, 대부분 메모리 유지를 로컬적인 의사 결정 문제로 취급하며, 현실적인 관찰 가능성 제약 조건 하에서의 장기적인 결과를 명시적으로 모델링하지 않습니다. 이러한 격차를 해소하기 위해, 우리는 메모리 유지를 명시적인 예산 적합성, 증거의 유용성 및 누락 페널티, 재획득 지연 및 오래된 정보 위험을 포함하는 지연 비용과 함께 제약 조건이 있는 확률적 최적화 문제로 공식화합니다. 그런 다음, 우리는 관찰 가능성을 고려한 안전한 학습(OSL-MR) 프레임워크를 제안하며, 이는 온라인에서 관찰 가능한 특징과 오프라인에서 사용 가능한 지도 데이터 간의 엄격한 분리(OAS)를 강제합니다. OSL-MR은 실현된 증거 기반으로 훈련된 증거 학습기와 혼합 점수(Mixed-Score) 휴리스틱을 결합합니다. 혼합 점수 휴리스틱은 배포 가능한 온라인 안전 기준 역할을 하며, 또한 학습을 위한 구조화된 귀납적 사전 지식을 제공합니다. 결과적으로 생성되는 정책은 상호 작용 데이터를 통해 쿼리 조건부 증거 값을 직접 학습하며, 동시에 동일한 관찰 가능성 제약 조건 하에서 배포될 수 있습니다. LOCOMO 및 LongMemEval에 대한 실험 결과는 OSL-MR이 최근 정보 기반 방법, Generative Agents 스타일의 점수 부여 및 기타 휴리스틱 기준보다 일관되게 우수한 성능을 보인다는 것을 보여줍니다. 특히 메모리 예산이 제한적인 경우 더욱 그렇습니다. 혼합 점수 사전 지식은 정밀도를 더욱 향상시키는 동시에 재현율을 유지하며, 민감도 분석은 다양한 비용 구성에서 견고성을 입증합니다.
Long-horizon language agents accumulate observations, reasoning traces, and retrieved facts that exceed their finite context windows, making memory retention a fundamental resource-allocation problem. Existing memory systems improve management through heuristic scoring, retrieval optimization, or learned compression, but largely treat retention as a local decision problem and do not explicitly model its long-term consequences under realistic observability constraints. To fill this gap, we formulate memory retention as a constrained stochastic optimization problem with explicit budget feasibility, evidence utility, and delayed costs including miss penalties, reacquisition delays, and stale-information risk. We then propose OSL-MR (Observability-Safe Learning for Memory Retention), a novel framework that enforces a strict separation between online-observable features and offline-available supervision (OAS). OSL-MR combines an evidence learner trained from realized evidence supervision with a Mixed-Score heuristic that serves both as a deployable online-safe baseline and as a structured inductive prior for learning. The resulting policy learns query-conditioned evidence value directly from interaction data while remaining deployable under the same observability constraints. Experiments on LOCOMO and LongMemEval show that OSL-MR consistently outperforms recency-based methods, Generative Agents-style scoring, and other heuristic baselines, particularly under tight memory budgets. The Mixed-Score prior further improves precision while preserving recall, and sensitivity analysis demonstrates robustness across a wide range of cost configurations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.