OpsMem: 교차 메모리 공명을 활용한 이중 메모리 추론을 통한 오류 진단
OpsMem: Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis
현대 소프트웨어 시스템의 오류 진단은 운영 경험에 기반하여 반복적인 증거 수집 및 가설 추론이 필요합니다. 기존 LLM 기반 방법들은 에이전트 추론 또는 지식 보강을 통해 진단을 개선하지만, 반복적인 진단 과정에서 진화하는 진단 상태와 운영 경험 간의 조율 메커니즘이 부족한 경우가 많습니다. 본 논문에서는 OpsMem이라는 이중 메모리 프레임워크를 제안합니다. OpsMem은 현재 진단 상태에 대한 단기 메모리와 재사용 가능한 운영 경험에 대한 장기 메모리를 유지합니다. OpsMem은 교차 메모리 공명을 사용하여 현재 상태와 관련된 장기 메모리를 활성화하고, 단기 메모리와 활성화된 장기 메모리를 기반으로 다중 에이전트 진단을 수행하며, 해결된 문제에서 얻은 재사용 가능한 경험을 장기 메모리에 통합합니다. 실제 Huawei 마이크로 서비스 오류 진단 데이터 세트에 대한 실험 결과, OpsMem은 대표적인 에이전트 추론 및 지식 보강 기반 모델들을 능가하는 성능을 보여주었으며, 각각 Match와 Relevant 지표에서 가장 강력한 기준 모델보다 최대 46.88% 및 18.39%의 성능 향상을 달성했습니다.
Failure diagnosis in modern software systems requires iterative evidence acquisition and hypothesis reasoning guided by operational experience. Existing LLM-based methods improve diagnosis through agentic reasoning or knowledge augmentation, but they often lack a mechanism to coordinate the evolving diagnostic state with operational experience during iterative diagnosis. We propose OpsMem, a dual-memory framework that maintains a short-term memory for the current diagnostic state and a long-term memory for reusable operational experience. OpsMem uses cross-memory resonance to activate state-relevant long-term memory, conditions multi-agent diagnosis on the short-term and activated long-term memories, and consolidates reusable experience from solved incidents back into long-term memory. Experiments on a real-world Huawei microservice failure diagnosis dataset show that OpsMem outperforms representative agentic-reasoning and knowledge-augmented baselines, improving Match and Relevant by up to 46.88% and 18.39% over the strongest baseline, respectively.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.