2606.10431v1 Jun 09, 2026 cs.CV

다중 작업 차량 경로 문제 해결을 위한 시각 기반 기초 모델

Vision-Assisted Foundation Model for Solving Multi-Task Vehicle Routing Problems

Zhiguang Cao
Zhiguang Cao
Citations: 10
h-index: 2
Y. Ong
Y. Ong
Citations: 10
h-index: 2
Shuangchun Gui
Shuangchun Gui
Citations: 84
h-index: 3
Wen Song
Wen Song
Citations: 3
h-index: 1

다중 작업 차량 경로 문제는 다양한 산업 및 서비스 분야에서 효율성을 향상시키는 데 중요한 역할을 합니다. 이러한 문제는 여러 변형으로 구성되며, 고객의 다양한 제약 조건을 충족하면서 경로 비용을 최적화합니다. 기존의 다중 작업 VRP 솔버는 그래프 기반 방식만 사용하며, 이는 여러 제약 조건이 있는 경우에도 적절하게 처리할 수 없다는 한계를 가지고 있습니다. 시각 정보는 복잡한 의미를 표현하는 방식으로 다양한 VRP 제약 조건을 인코딩하는 데 큰 잠재력을 보여줍니다. 본 연구에서는 이러한 점에 착안하여 시각 이미지에서 패치 수준의 의미 정보를 학습하고, 이를 그래프 기반 모델과 통합하여 다양한 VRP 변형을 동시에 해결하고자 합니다. 그러나 이 접근 방식을 다중 작업 VRP에 직접 적용하는 것은 다음과 같은 세 가지 과제를 안고 있습니다. 1) 기존의 VRP 이미지는 다중 작업 VRP에 필수적인 제약 조건 표현이 부족합니다. 2) 개별 패치의 고정된 수용 영역은 다양한 작업 간의 요구 사항을 효과적으로 충족하지 못할 수 있습니다. 3) 제약 조건 간의 불균형한 픽셀 분포는 모델이 픽셀 수가 적은 제약 조건을 간과하게 만들 수 있습니다. 본 논문에서는 이러한 과제를 해결하기 위해 시각 기반 기초 모델(VaFM)을 제안합니다. 시각 모달리티에서는 컨볼루션 신경망을 사용하여 모든 제약 조건에 맞는 이미지를 인코딩합니다. 얻어진 패치 임베딩은 그래프 기반 노드와 융합되어 솔루션을 생성하며, 픽셀 불균형 문제를 해결하기 위한 보조 작업이 설계되었습니다. VaFM의 성능은 16가지 다양한 VRP 변형을 통해 평가되었으며, 실험 결과는 VaFM이 최첨단 방법보다 우수하며, 특히 복잡한 제약 조건을 가진 경우에 더욱 그렇다는 것을 보여줍니다.

Original Abstract

Multi-task vehicle routing problems play a critical role in enhancing efficiency across various industries and service sectors. These problems consist of multiple variants that optimize routing costs while meeting diverse customer constraints. Existing multi-task VRP solvers solely utilize a graph-based modality, limiting their ability to address variants with multiple constraints. As a format to represent complex semantics, vision modality shows great potential for encoding diverse VRP constraints. This motivates us to learn patch-level semantics from the vision images, and then integrate them into a graph-based model to solve various VRP variants simultaneously. However, directly applying this approach to multi-task VRPs presents three challenges: 1) existing VRP images lack constraint representations, which are essential for multi-task VRPs, 2) the fixed receptive field of individual patches cannot effectively accommodate varying requirements across tasks, and 3) imbalanced pixel distribution among constraints may cause the model to overlook constraints with fewer pixels. In this paper, we propose a vision-assisted foundation model (VaFM) to address these challenges. In the vision modality, input images tailored to all constraints are encoded by a convolutional neural network. The obtained patch embeddings are fused with graph-based nodes to generate solutions, with an auxiliary task designed to address the pixel-imbalanced issue. The performance of VaFM is evaluated across 16 different VRP variants. The experimental results demonstrate the superiority of VaFM over state-of-the-art methods, especially for variants with complex constraints.

1 Citations
0 Influential
1.5 Altmetric
8.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!