비전 언어 모델을 활용한 산업 현장 문제 해결 지침에서 절차적 지식 추출
Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models
산업 현장 문제 해결 지침은 흐름도와 유사한 다이어그램 형태로 작성되며, 공간적 구성과 기술적인 언어가 결합되어 의미를 전달합니다. 이러한 지식을 작업 현장 인력의 장비 문제 진단 및 해결을 지원하는 시스템에 통합하기 위해서는 먼저 정보를 추출하고 기계가 해석할 수 있도록 구조화해야 합니다. 그러나 수동으로 정보를 추출하는 과정은 노동 집약적이고 오류가 발생하기 쉽습니다. 비전 언어 모델은 시각적 및 텍스트적 의미를 동시에 해석하여 이 프로세스를 자동화할 수 있는 잠재력을 가지고 있지만, 이러한 모델의 성능은 해당 지침에 대한 분석이 아직 부족합니다. 본 논문에서는 구조화된 지식 추출을 위해 두 가지 비전 언어 모델을 평가하고, 표준적인 지시 기반 프롬프트와 문제 해결 레이아웃 패턴을 활용한 강화된 접근 방식의 두 가지 프롬프트 전략을 비교합니다. 결과는 모델별로 레이아웃 민감도와 의미론적 견고성 간의 균형이 다르다는 것을 보여주며, 실제 적용 시 고려해야 할 사항을 제시합니다.
Industrial troubleshooting guides encode diagnostic procedures in flowchart-like diagrams where spatial layout and technical language jointly convey meaning. To integrate this knowledge into operator support systems, which assist shop-floor personnel in diagnosing and resolving equipment issues, the information must first be extracted and structured for machine interpretation. However, when performed manually, this extraction is labor-intensive and error-prone. Vision Language Models offer potential to automate this process by jointly interpreting visual and textual meaning, yet their performance on such guides remains underexplored. This paper evaluates two VLMs on extracting structured knowledge, comparing two prompting strategies: standard instruction-guided versus an augmented approach that cues troubleshooting layout patterns. Results reveal model-specific trade-offs between layout sensitivity and semantic robustness, informing practical deployment decisions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.