실행에서 능력으로: 절차적 지식 합성 기반 과학적 경험 통합
From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis
최근 대규모 언어 모델은 과학 컴퓨팅 작업을 점점 더 많이 수행하지만, 하나의 문제에서 얻은 실행 가능한 피드백이 후속 문제에 대한 지속적인 역량으로 이어지는 경우는 드뭅니다. 본 연구에서는 과학 컴퓨팅 경험 통합을 탐구합니다. 즉, 검증된 런타임 경험을 전송 가능한 절차적 지식과 지속적인 모델 개선으로 변환하는 것입니다. 이 설정은 두 가지 과제를 제시합니다. 첫째, 경로에서 파생된 아티팩트는 특정 소스에 대한 수정 사항을 인코딩할 수 있으며, 다양한 작업에 적용되는 계산 메커니즘을 반영하지 않을 수 있습니다. 둘째, 약한 대상 모델은 유효한 추상 절차를 실제로 실행하는 데 어려움을 겪을 수 있는데, 이는 추상화와 실행 간의 격차로 인해 발생합니다. 우리는 SciConsolidate라는 방법을 제안합니다. 이 방법은 검증된 성공 및 실패 사례를 비교하여 작업 간 절차를 유도하고, 개발-검증 단계를 통해 이를 선택하며, 기존의 참조 답변 없이 통합 데이터를 확장하기 위해 실패 정보를 활용한, 답이 없는 질의 생성(answer-free query synthesis)을 사용합니다. 대상 모델이 이러한 추상화를 직접 실행하지 못할 수 있으므로, 더 강력한 모델은 이를 표준적인 절차 기반 학습(SFT)을 위한 실행 가능한 코드로 구체화합니다. 또한, 절차가 없는 교사 모델 브랜치를 사용하여 절차적 안내의 가치를 분리합니다. SciCode 데이터셋에서 런타임 절차 주입은 Qwen3.6-27B 모델의 성능을 +3.85/+6.26(부분 단계/전체 문제)만큼 향상시켰지만, Qwen3.5-9B 모델에서는 거의 개선 효과가 없었습니다. 이는 추상화와 실행 간의 격차에 대한 실질적인 증거를 제공합니다. 절차 기반 구체화를 통해 9B 모델은 절차가 없는 학습 환경에서 SFT 제어 모델보다 +3.89/+6.25만큼, 원래 9B 모델보다 +5.62/+11.25만큼 성능이 향상되었습니다. 이러한 결과는 과학 컴퓨팅을 위한 경험-능력 전환 경로를 제시하며, 자기 개선형 과학 지원 시스템의 확대를 위한 실질적인 기반을 제공합니다.
Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capability on subsequent problems. We study scientific-computing experience consolidation: converting verified runtime experience into transferable procedural knowledge and persistent model improvement. This setting presents two challenges: trajectory-derived artifacts may encode source-specific repairs rather than cross-task computational mechanisms; and a weaker target model may be unable to operationalize an otherwise valid abstract procedure - an abstraction-execution gap. We introduce SciConsolidate, which contrasts verified successes and failures to induce cross-task procedures, selects them through a development-validation gate, and uses failure-informed, answer-free query synthesis to expand the consolidation data without requiring pre-existing reference answers. Because the target model may not directly execute these abstractions, a stronger model concretizes them into executable code supervision for standard, procedure-free SFT; a matched no-procedure teacher branch isolates the value of procedural guidance. On SciCode, runtime procedure injection improves Qwen3.6-27B by +3.85/+6.26 sub-step/main-problem points, but yields almost no aggregate main-problem gain for Qwen3.5-9B, providing operational evidence of the abstraction-execution gap. After procedure-guided concretization, the 9B student improves under procedure-free deployment by +3.89/+6.25 points over the no-procedure SFT control and by +5.62/+11.25 over the original 9B model. These results establish an experience-to-capability pathway for scientific computing and provide a practical starting point for scaling self-improving scientific assistance.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.