PPVC 모듈 공장 내 유연한 작업장 스케줄링을 위한 시간 지연 인지 심층 강화 학습
Time-Lag-Aware Deep Reinforcement Learning for Flexible Job-Shop Scheduling in PPVC Module Factories
프리패브릭 프리피니시드 볼륨 건설(PPVC)은 대부분의 건축 작업을 모듈 공장에 이전하며, 해당 공장의 생산 현장은 유연한 작업장으로 운영됩니다. 주요 문제점은 콘크리트 경화, 방수 테스트 및 페인트 건조로 인해 발생하는 긴 후공정 시간 지연입니다. 이로 인해 모듈이 특정 워크스테이션에 묶이는 동안 해당 워크스테이션은 비어 있게 됩니다. 공식 국가 프리패브릭 가이드라인을 기반으로 한 표준 데이터셋에서, 이러한 지연은 최적의 기준 처리 시간을 평균적으로 약 67% 증가시킵니다. 의사 결정 시 이러한 지연을 무시하고 사후적으로 수정하는 방법은 모든 작업 규칙보다 성능이 떨어집니다. 본 연구에서는 시간 지연을 고려한 동역학 모델과 허용 가능한 보상 범위, 예측적인 시간 지연 특징 채널, 그리고 활성 상태를 마스크 처리한 작업 및 스테이션 유형 임베딩이라는 세 가지 최소 침습적이고 독립적으로 검증 가능한 확장 기능을 통해 최첨단 듀얼 어텐션 심층 강화 학습 솔버를 개선했습니다. 각 확정 기능이 비활성화되면 구현은 원래 솔버와 정확히 동일한 결과를 나타내므로, 모든 성능 향상은 이러한 수정 사항에 기인합니다. 본 연구에서는 공식 가이드라인을 기반으로 한 공개 데이터셋 생성기를 제공합니다. 검증되지 않은 데이터셋에서 학습된 정책은 가장 강력한 솔버-프리 스케줄러이며, 제약 조건 프로그래밍 기준의 약 4% 이내의 성능을 보이며 모든 작업 규칙과 유전 알고리즘 메타 휴리스틱을 능가합니다. 이러한 우위는 생산 능력 부족 상황에서 더욱 두드러지며, 단일 크기 혼합 정책이 학습된 공장 규모 범위 전체에 걸쳐 이러한 우위를 유지합니다. 본 시스템은 솔버, 모델 또는 라이선스가 필요 없으며, 문제 발생 후 몇 초 내에 재계획이 가능합니다. 정확한 솔버를 사용할 수 있는 경우에도, 해당 솔버는 성능의 상한선을 나타내며, 우리는 이러한 경계를 명시적으로 정의했습니다.
Prefabricated prefinished volumetric construction moves most building work into module factories, whose production floor operates as a flexible job shop. A major complication is decisive: long post-operation time-lags caused by concrete curing, watertightness ponding tests, and paint drying, during which a module is blocked while its workstation stays free. On benchmark instances grounded in an official national prefabrication guidebook, these lags inflate even the optimal reference makespan by about 67% on average, and ignoring them at decision time, then repairing to feasibility, is worse than every dispatching rule. We adapt a state-of-the-art dual-attention deep reinforcement learning solver through three minimally invasive, individually ablatable extensions: lag-aware dynamics with an admissible reward bound, two anticipatory lag feature channels, and liveness-masked operation- and station-type embeddings. With every extension disabled the implementation reproduces the original solver exactly, so all gains are attributable to the adaptations. We release a public, guidebook-grounded benchmark generator. On held-out instances the learned policy is the strongest solver-free scheduler: it reaches within about 4% of a constraint-programming reference and beats every dispatching rule and a genetic-algorithm metaheuristic, with its advantage widening under capacity contention, and a single size-mixed policy carries this lead across the trained range of factory sizes. It needs no solver, model, or license in the loop and re-plans within seconds of a disruption; where an exact solver can be deployed, that solver remains the quality ceiling, a boundary we map explicitly.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.