2604.17966v1 Apr 20, 2026 cs.AI

TPS-CalcBench: 극초음속 열 보호 시스템 엔지니어링 분야에서 LLM의 분석 계산 능력 평가 및 진단 프레임워크

TPS-CalcBench: A Benchmark and Diagnostic Evaluation Framework for LLM Analytical Calculation Competence in Hypersonic Thermal Protection System Engineering

Jinglai Zheng
Jinglai Zheng
Citations: 87
h-index: 5
Haiming Huang
Haiming Huang
Citations: 71
h-index: 4
Chuhang Qiao
Chuhang Qiao
Citations: 0
h-index: 0

안전이 중요한 항공우주 엔지니어링 분야에서 LLM을 추론 보조 도구로 활용하기 위해서는 일반적인 과학 벤치마크보다 엄격한 평가 기준이 필요합니다. 특히, 극초음속 열 보호 시스템(TPS) 설계에서 부정확한 정지점 열유속 또는 경계층 계산은 심각한 설계 오류를 초래할 수 있습니다. 수치적으로는 합리적이지만 물리적으로는 잘못된 결과를 제공하는 모델은 응답을 거부하는 모델보다 더 위험합니다. 현재의 과학 벤치마크는 추상적인 수학 및 기초 물리학만 테스트하고, 최종 결과만을 평가하며, 공학적 추론 과정을 고려하지 않아 이러한 중요한 오류를 감지할 수 없습니다. 우리는 극초음속 공기역학과 고온 가스 역학 분야에서 경험이 풍부한 TPS 엔지니어가 시뮬레이션 없이 수행하는 닫힌 형태의 분석 계산을 위한 첫 번째 진단 벤치마크인 TPS-CalcBench를 제안합니다. 우리의 주요 기여 내용은 다음과 같습니다. Anderson 교재를 기반으로 4가지 난이도 수준과 8가지 범주로 구성된 도메인 특화된 작업 분류, 8가지 평가 기준을 통해 결과 정확도와 추론 품질을 측정하는 이중 추적 평가, 인간 전문가의 검토를 통해 정확한 답과 잘못된 추론을 식별하는 교정된 평가 시스템, 인간-AI 데이터 파이프라인을 통해 4560개의 원시 데이터에서 420개의 높은 신뢰도를 가진 핵심 항목과 810개의 노이즈 제어된 사전 필터링 항목을 생성, 모델 순위에 미치는 데이터 품질 영향을 측정하는 노이즈 민감도 분석, 그리고 DFA-TPS 미세 조정, RAG-EQ 검색 기반 근거, PA-CoT 프로세스 인식 프롬프트의 세 가지 진단 개입 방법입니다. 7개 그룹에서 개발된 13개의 모델을 대상으로 실시한 테스트 결과, 성능 차이가 매우 컸으며(KPI 12.6-87.9), 숨겨진 공식 선택 오류, 데이터 기반 순위 변화, 그리고 효과적인 개입 개선 효과가 확인되었습니다. 이러한 결과를 바탕으로 안전이 중요한 공학 분야의 LLM 배포 평가를 위한 완전한 진단-평가-개입 프레임워크를 구축했습니다.

Original Abstract

Deploying LLMs as reasoning assistants in safety-critical aerospace engineering requires stricter evaluation criteria than general scientific benchmarks. In hypersonic thermal protection system (TPS) design, inaccurate stagnation-point heat flux or boundary-layer calculations may cause catastrophic design margin violations. Models with numerically reasonable but physically invalid answers are more dangerous than those declining to respond. Current scientific benchmarks only test abstract math and basic physics, evaluate final answers solely, ignore engineering reasoning processes, and cannot detect such critical failures. We propose TPS-CalcBench, the first diagnostic benchmark for closed-form analytical calculations in hypersonic aerodynamics and high-temperature gas dynamics that experienced TPS engineers conduct without simulations. Our contributions include domain-oriented task taxonomy with 4 difficulty levels and 8 categories from Anderson's textbook, dual-track evaluation measuring result accuracy and reasoning quality via an 8-dimension rubric and calibrated judge with human audit to identify right answer wrong reasoning issues, human-AI data pipeline producing 420 high-confidence core items and 810 noise-controlled pre-gating items from 4560 raw data, noise-sensitivity analysis measuring data quality impacts on model ranking, and three diagnostic intervention methods: DFA-TPS fine-tuning, RAG-EQ retrieval grounding and PA-CoT process-aware prompting. Tests on 13 models from 7 groups show wide performance differences (KPI 12.6-87.9), hidden formula selection defects, data-driven rank changes and effective intervention improvements, establishing a complete diagnose-evaluate-intervene framework for safety-critical engineering LLM deployment assessment.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!