조립 라인 장애 복구 시스템에서 단계별 가이드 주입을 활용한 순환형 MAPPO
Phase-Aware Guidance Injection for Recurrent MAPPO in Assembly-Line Disruption Recovery
산업용 조립 라인의 장애 복구는 기계 고장, 작업자 부재 및 긴급 주문 상황에서 신속한 의사 결정을 필요로 합니다. 기존 방법들은 경직된 수동으로 설계된 복구 로직에 의존하거나, 비정상적인 복구 시간(ART)을 줄이고 정시 납품률(OTD)을 유지하기 위해 의사 결정 시점에 다양한 외부 복구 지식을 효과적으로 활용하지 못하는 적응형 정책을 학습합니다. 이러한 문제를 해결하기 위해, 우리는 훈련된 순환형 MAPPO(RMAPPO) 스케줄링 정책에 평가 단계에서 로짓 레벨의 행동 편향을 통해 가이드 정보를 주입하는 단계별 가이드 주입 프레임워크를 제안합니다. 이 프레임워크는 규칙 기반, 리플레이 기반 및 온라인 LLM 기반 가이드를 위한 통합적인 의사 결정 인터페이스를 제공하며, 비정상 상태 및 복구 단계에서만 개입을 활성화합니다. 사용자 정의된 AssemblyLineEnv 환경에서의 실험 결과, 고품질의 규칙 기반 가이드가 가장 큰 성능 향상을 가져왔으며, 리플레이 기반 가이드는 불완전한 사용 가능성 하에서도 안정적으로 작동하고, 온라인 LLM 기반 가이드는 여전히 유용한 중간 수준의 개선을 제공합니다. 이러한 결과는 의사 결정 시점의 가이드 주입이 액터 네트워크를 재설계하지 않고도 다양한 복구 힌트를 활용할 수 있음을 보여줍니다.
Disruption recovery in industrial assembly lines requires timely decisions under machine faults, worker absence, and emergency orders. Existing methods either rely on rigid handcrafted recovery logic or learn adaptive policies that do not readily exploit heterogeneous external recovery knowledge at decision time to reduce abnormal recovery time (ART) and preserve on-time delivery (OTD). To address this gap, we propose a phase-aware guidance injection framework that augments a trained recurrent MAPPO (RMAPPO) scheduling policy through logit-level action bias during evaluation. The framework provides a unified decision-time interface for rule-based, replay-based, and online LLM-based guidance, while activating intervention only during abnormal and recovery phases. Experiments on a custom AssemblyLineEnv show that high-quality rule guidance yields the strongest gains, replay-based guidance degrades smoothly under imperfect availability, and online LLM guidance still provides useful intermediate improvements. These results show that decision-time guidance injection can exploit heterogeneous recovery hints without redesigning the actor.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.