복잡 동적 시스템에서의 피드백 제어기 설계에 대한 대규모 언어 모델의 성능 평가 및 지식 증류
Benchmarking and Reasoning Distillation of Large Language Models for Feedback Controller Design in Complex Dynamical Systems
대규모 언어 모델(LLM)은 다양한 과학 분야에서 놀라운 능력을 보여주었지만, 피드백 제어기 설계는 아직 충분히 연구되지 않았습니다. 기존의 벤치마크는 주로 선형 단일 자유도(DoF) 시스템과 대규모 API 기반 모델에 초점을 맞추고 있으며, 복잡한 제어기 설계 작업에서의 성능과 에지 환경 배포 가능성에 대한 명확성이 부족합니다. 이러한 한계를 극복하기 위해, 우리는 대규모 언어 모델을 위한 복잡 동역학-제어 벤치마크(CoDyControlBench)를 소개합니다. 이 벤치마크는 자유도 수, 시스템 유형, 결합 수준, 감쇠 모드 및 제어기 유형의 다섯 가지 평가 차원을 포함하여 총 132개의 시스템 구성으로 이루어져 있습니다. GPT, Gemini, Claude와 같은 세 개의 상용 모델과 GLM, DeepSeek, Qwen과 같은 세 개의 오픈 소스 모델을 포함한 최첨단 LLM 6개를 독립적인 세 번의 실행에 걸쳐 평가했습니다. GPT는 94.8%로 가장 높은 설계 성공률을 달성했지만, Qwen은 50.0%로 가장 낮은 성공률을 보였습니다. 벤치마크 차원별로 자유도와 제어기 유형이 모델 평균적으로 가장 큰 설계 성공률 변화를 나타냈으며, 각각 36.3% 및 17.6%의 범위를 보였는데, 이는 시스템 유형, 결합 수준 및 감쇠 모드와 관련된 범위보다 높습니다. GPT와 Qwen의 성능 차이는 주로 제어 설계 지식, 특히 이득 선택과 과도 현상을 제한하는 메커니즘 사용에서 비롯되었습니다. 에지 환경 배포를 위해 15억 개의 파라미터를 가진 특수 모델을 지식 증류 방식으로 개발했습니다. 지식 증류된 모델은 CoDyControlBench에서 답변 증류 모델 및 기본 모델보다 우수한 성능을 보였으며, 1~6개의 자유도 범위에서 안정적인 성능을 유지했으며, 공압 인공 근육 구동 로봇 팔에 대한 세 가지 물리적 실험 모두에서 목표 추적에 성공했습니다. 이러한 결과는 벤치마크 기준점을 제시하고 경량의 에지 배포 가능 제어기 설계 모델의 잠재력을 강조합니다.
Although remarkable capabilities have been demonstrated by Large Language Models (LLMs) across scientific domains, feedback controller design remains underexplored. Existing benchmarks focus mainly on linear single-Degree-of-Freedom (DoF) systems and large API-hosted models, leaving performance on complex controller-design tasks and feasibility for edge deployment unclear. To address these limitations, we introduce the Complex Dynamics-to-Control Benchmark for Large Language Models (CoDyControlBench), comprising 132 system configurations across five evaluation dimensions: number of DoF, system type, coupling level, damping regime, and controller type. Six state-of-the-art LLMs were evaluated over three independent runs, including three commercial models (GPT, Gemini, and Claude) and three open-source models (GLM, DeepSeek, and Qwen). GPT achieved the highest design success rate at 94.8\%, whereas Qwen showed the lowest rate at 50.0\%. Across the benchmark dimensions, DoF and controller type exhibited the largest model-averaged variations in design success, with success-rate ranges of 36.3\% and 17.6\%, respectively, exceeding those associated with system type, coupling level, and damping regime. Comparison of GPT and Qwen showed that their performance gap arose mainly from the control-design knowledge, particularly gain selection and the use of transient-limiting mechanisms. For edge deployment, a specialized 1.5B-parameter model was developed through reasoning distillation. The reasoning-distilled model outperformed the answer-distilled and base model on CoDyControlBench, maintained stable performance across 1-6 DoFs, and achieved successful traget tracking in all three physical trials on a pneumatic-artificial-muscle-driven robotic arm. These results establish a benchmark baseline and highlight the potential of lightweight, edge-deployable controller-design models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.