ShipTraj-R1: 그룹 상대 정책 최적화(Group Relative Policy Optimization)를 활용한 대규모 언어 모델 기반 선박 항로 예측 강화
ShipTraj-R1: Reinforcing Ship Trajectory Prediction in Large Language Models via Group Relative Policy Optimization
최근 강화 학습 기반 미세 조정 기술은 대규모 언어 모델(LLM)의 추론 능력을 크게 향상시켰습니다. 특히, 그룹 상대 정책 최적화(GRPO)와 같은 방법은 다양한 분야에서 뛰어난 성능을 보여주었습니다. 그러나 LLM을 선박 항로 예측에 적용하는 연구는 아직 미흡한 실정입니다. 본 논문에서는 선박 항로 예측을 텍스트-텍스트 생성 문제로 재구성하는 새로운 LLM 기반 프레임워크인 ShipTraj-R1을 제안합니다. (1) 모델이 적응적인 사고 과정(Chain-of-Thought, CoT)을 수행하도록 유도하기 위해, 충돌 가능성이 있는 선박의 항로 정보를 포함하는 동적 프롬프트를 설계했습니다. (2) 모델의 추론 형식과 예측 정확도를 장려하기 위해, 포괄적인 규칙 기반 보상 메커니즘을 도입했습니다. (3) 제안하는 ShipTraj-R1은 도메인 특화 프롬프트와 보상을 통해 GRPO 메커니즘으로 강화되며, Qwen3을 모델의 기반으로 사용합니다. 두 가지 복잡하고 실제 해양 데이터셋에 대한 광범위한 실험 결과는 제안하는 ShipTraj-R1이 최첨단 딥 러닝 및 LLM 기반 모델과 비교하여 가장 낮은 오차율을 달성함을 보여줍니다.
Recent advancements in reinforcement fine-tuning have significantly improved the reasoning ability of large language models (LLMs). In particular, methods such as group relative policy optimization (GRPO) have demonstrated strong capabilities across various fields. However, applying LLMs to ship trajectory prediction remains largely unexplored. In this paper, we propose ShipTraj-R1, a novel LLM-based framework that reformulates ship trajectory prediction as a text-to-text generation problem. (1) We design a dynamic prompt containing trajectory information about conflicting ships to guide the model to achieve adaptive chain-of-thought (CoT) reasoning. (2) We introduce a comprehensive rule-based reward mechanism to incentivize the reasoning format and prediction accuracy of the model. (3) Our ShipTraj-R1 is reinforced through the GRPO mechanism guided by domain-specific prompts and rewards, and utilizes the Qwen3 as the model backbone. Extensive experimental results on two complex and real-world maritime datasets show that the proposed ShipTraj-R1 achieves the least error compared with state-of-the-art deep learning and LLM-based baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.