2607.24459v1 Jul 27, 2026 cs.AI

실행에서 능력으로: 절차적 지식 합성 기반 과학적 경험 통합

From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis

Jiahao Zhao
Jiahao Zhao
Citations: 424
h-index: 8
Liwei Dong
Liwei Dong
Citations: 24
h-index: 2
Nan Xu
Nan Xu
Citations: 2
h-index: 1

최근 대규모 언어 모델은 과학 컴퓨팅 작업을 점점 더 많이 수행하지만, 하나의 문제에서 얻은 실행 가능한 피드백이 후속 문제에 대한 지속적인 역량으로 이어지는 경우는 드뭅니다. 본 연구에서는 과학 컴퓨팅 경험 통합을 탐구합니다. 즉, 검증된 런타임 경험을 전송 가능한 절차적 지식과 지속적인 모델 개선으로 변환하는 것입니다. 이 설정은 두 가지 과제를 제시합니다. 첫째, 경로에서 파생된 아티팩트는 특정 소스에 대한 수정 사항을 인코딩할 수 있으며, 다양한 작업에 적용되는 계산 메커니즘을 반영하지 않을 수 있습니다. 둘째, 약한 대상 모델은 유효한 추상 절차를 실제로 실행하는 데 어려움을 겪을 수 있는데, 이는 추상화와 실행 간의 격차로 인해 발생합니다. 우리는 SciConsolidate라는 방법을 제안합니다. 이 방법은 검증된 성공 및 실패 사례를 비교하여 작업 간 절차를 유도하고, 개발-검증 단계를 통해 이를 선택하며, 기존의 참조 답변 없이 통합 데이터를 확장하기 위해 실패 정보를 활용한, 답이 없는 질의 생성(answer-free query synthesis)을 사용합니다. 대상 모델이 이러한 추상화를 직접 실행하지 못할 수 있으므로, 더 강력한 모델은 이를 표준적인 절차 기반 학습(SFT)을 위한 실행 가능한 코드로 구체화합니다. 또한, 절차가 없는 교사 모델 브랜치를 사용하여 절차적 안내의 가치를 분리합니다. SciCode 데이터셋에서 런타임 절차 주입은 Qwen3.6-27B 모델의 성능을 +3.85/+6.26(부분 단계/전체 문제)만큼 향상시켰지만, Qwen3.5-9B 모델에서는 거의 개선 효과가 없었습니다. 이는 추상화와 실행 간의 격차에 대한 실질적인 증거를 제공합니다. 절차 기반 구체화를 통해 9B 모델은 절차가 없는 학습 환경에서 SFT 제어 모델보다 +3.89/+6.25만큼, 원래 9B 모델보다 +5.62/+11.25만큼 성능이 향상되었습니다. 이러한 결과는 과학 컴퓨팅을 위한 경험-능력 전환 경로를 제시하며, 자기 개선형 과학 지원 시스템의 확대를 위한 실질적인 기반을 제공합니다.

Original Abstract

Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capability on subsequent problems. We study scientific-computing experience consolidation: converting verified runtime experience into transferable procedural knowledge and persistent model improvement. This setting presents two challenges: trajectory-derived artifacts may encode source-specific repairs rather than cross-task computational mechanisms; and a weaker target model may be unable to operationalize an otherwise valid abstract procedure - an abstraction-execution gap. We introduce SciConsolidate, which contrasts verified successes and failures to induce cross-task procedures, selects them through a development-validation gate, and uses failure-informed, answer-free query synthesis to expand the consolidation data without requiring pre-existing reference answers. Because the target model may not directly execute these abstractions, a stronger model concretizes them into executable code supervision for standard, procedure-free SFT; a matched no-procedure teacher branch isolates the value of procedural guidance. On SciCode, runtime procedure injection improves Qwen3.6-27B by +3.85/+6.26 sub-step/main-problem points, but yields almost no aggregate main-problem gain for Qwen3.5-9B, providing operational evidence of the abstraction-execution gap. After procedure-guided concretization, the 9B student improves under procedure-free deployment by +3.89/+6.25 points over the no-procedure SFT control and by +5.62/+11.25 over the original 9B model. These results establish an experience-to-capability pathway for scientific computing and provide a practical starting point for scaling self-improving scientific assistance.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!