2607.01595v1 Jul 02, 2026 cs.AI

안전하고 적응적인 클라우드 자가 복구: 신경-기호 세계 모델을 활용한 LLM 생성 복구 계획 검증

Safe and Adaptive Cloud Healing: Verifying LLM-Generated Recovery Plans with a Neural-Symbolic World Model

Junyan Tan
Junyan Tan
Citations: 0
h-index: 0
Yichen Fang
Yichen Fang
Citations: 0
h-index: 0
Xinyue Luo
Xinyue Luo
Citations: 0
h-index: 0
Tian Shen
Tian Shen
Citations: 0
h-index: 0
Zeyu Qiao
Zeyu Qiao
Citations: 0
h-index: 0
Haoran Lin
Haoran Lin
Citations: 0
h-index: 0
Siyuan Guo
Siyuan Guo
Citations: 0
h-index: 0

클라우드 기반 AI 시스템의 규모와 복잡성이 계속 증가함에 따라, 빠른 오류 감지 및 적응적 복구를 통한 서비스 안정성 확보는 중요한 과제로 부상했습니다. 기존 방법들은 대규모 언어 모델(LLM)을 사용하여 의미론적 이해를 수행하고, 심층 강화 학습(DRL)을 통해 정책 최적화를 수행하지만, 종종 순차적이고 느슨하게 결합된 아키텍처에 의존하여 LLM의 생성 및 추론 능력을 충분히 활용하지 못합니다. 본 논문에서는 PASE라는 계획 인식 의미 자가 복구 엔진을 제안하며, 이는 복구를 신경-기호 프로그램 합성 작업으로 재정의하는 새로운 오류 자가 복구 프레임워크입니다. PASE는 LLM을 핵심 계획 합성 엔진으로 사용하여 의미 기반 기본 요소 라이브러리로부터 구조화된 복구 계획을 생성합니다. 신경-기호 세계 모델은 시뮬레이션을 통해 계획의 실현 가능성을 검증하고, DRL을 통해 학습된 메타 프롬프트 최적화기는 LLM의 계획 프로세스를 안내하는 최적의 프롬프트를 생성합니다. 이러한 밀접하게 결합된 추론-계획-검증-적응 루프는 미리 정의된 동작 공간을 넘어 동적이고 상황 인지적인 복구 전략 생성을 가능하게 합니다. 실제 클라우드 오류 주입 데이터 세트에 대한 실험 결과, PASE가 최첨단 방법보다 성능이 우수하며, 평균 시스템 복구 시간을 40% 이상 단축하고 알려지지 않은 오류 시나리오에서의 오류 감지 정확도를 향상시킴을 보여줍니다. 본 프레임워크는 LLM 기반 추론을 모델 지원 검증 및 메타 학습된 지침과 통합하여 자율 시스템 관리를 발전시킵니다.

Original Abstract

As the scale and complexity of cloud-based AI systems continue to escalate, ensuring service reliability through rapid fault detection and adaptive recovery has become a critical challenge. While existing approaches integrate Large Language Models (LLMs) for semantic understanding and Deep Reinforcement Learning (DRL) for policy optimization, they often rely on sequential, loosely coupled architectures that underutilize the generative and reasoning capabilities of LLMs. In this paper, we propose a paradigm shift with PASE, a Planning-Aware Semantic self-healing engine, a novel fault self-healing framework that reconceptualizes recovery as a neuro-symbolic program synthesis task. PASE employs an LLM as a core Plan Synthesis Engine to generate structured recovery plans from a library of semantic primitives. A Neural-Symbolic World Model verifies plan feasibility through simulation, while a Meta-Prompt Optimizer, trained via DRL, learns to generate optimal prompts that guide the LLM's planning process. This tight reason-plan-verify-adapt loop enables dynamic, context-aware recovery strategy generation beyond predefined action spaces. Experiments on a real-world cloud fault injection dataset demonstrate that PASE significantly outperforms state-of-the-art methods, reducing average system recovery time by over 40% and improving fault detection accuracy in unknown fault scenarios. Our framework advances autonomous system management by unifying LLM-based reasoning with model-assisted verification and meta-learned guidance.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!