2607.20058v1 Jul 22, 2026 cs.AI

오픈 소스 언어 모델에서 재료 과학 메커니즘 표현의 이해 및 제어

Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

Markus J. Buehler
Markus J. Buehler
Citations: 109
h-index: 4

대규모 언어 모델은 과학적 질문에 답변할 수 있지만, 정확한 답변이 모델이 관련 물리 법칙을 실제로 이해하고 사용하는지 나타내는 것은 아닙니다. 본 연구에서는 오픈 소스인 google/gemma-4-E4B-it 모델 내의 재료 과학 메커니즘 정보가 실험적으로 분리 가능한 세 가지 형태로 존재함을 보여줍니다. 즉, 개념은 개별 은닉 상태에서 읽을 수 있으며, 구성 요소는 상태 간의 제어된 변환에 의해 표현되며, 선택된 내부 표현은 공학적 답변에 인과적으로 영향을 미칩니다. 우리는 직접적인 단어 임베딩 및 야코비안 추출, 옵션 없는 상태 기하학, 60개의 반사실적 벤치마크 및 인과적 개입을 결합했습니다. 50개의 숨겨진 재료 설명에 대해 세 개의 독립적으로 학습된 야코비안 필터를 사용하여 개념 순위를 재현했으며, 두 가지 추출 방식에서 얻은 목표-무관 단어 집합을 통해 10개 중 9개의 메커니즘 분류를 정확하게 식별했습니다. 별도의 72개 질문 벤치마크는 메커니즘별 은닉 상태 영역을 생성했지만, 자세한 그래프 분석 결과 이러한 물리적 구성은 수치 비교로도 동일하게 설명될 수 있음을 확인했습니다. 따라서 우리는 물리적 입력 방향만 반대로 바꾼 동일한 프롬프트를 사용하여 숨겨진 상태의 움직임이 제공된 구성 법칙을 따르는지 여부를 확인했습니다. 이러한 상태 변환은 60개의 고정된 관계에서 직접적인, 물리적으로 중립적인 및 역방향 법칙을 순서대로 정렬했으며, 40개의 방향성 법칙 중 39개를 정확하게 분류했습니다. 어휘 기반 제어는 무작위 수준에 가까웠습니다. 양방향 개입은 모든 12개의 일치된 경우에서 답변 확률을 물리적으로 적절한 결과로 이동시키거나 그 반대로 만들었습니다. 또한, 반사실적 상태 패치는 메커니즘 및 답변 형식 간에 서로 반대되는 의사 결정 신호를 전송했습니다. 따라서 절대적인 상태보다 제어된 상태 변화를 통해 물리적 관계가 더 명확하게 드러났습니다.

Original Abstract

Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentally separable forms: concepts are readable in individual hidden states, constitutive orientation is carried by controlled transformations between states, and selected internal representations causally control engineering answers. We combine matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark and causal interventions. In 50 held-out materials descriptions, three independently fitted Jacobian lenses reproduced concept ranks, and target-free word sets from both readouts enabled blinded identification of 9 of 10 mechanism families. A separate 72-prompt benchmark produced mechanism-specific hidden-state neighborhoods, but an exact graph audit showed that this apparent physical organization was equally explained by numerical comparison. We therefore compared otherwise identical prompts in which only the direction of the physical input was reversed, asking whether the resulting hidden-state movement followed the supplied constitutive law. These state transformations ordered direct, physically neutral and inverse laws across 60 frozen relations and correctly oriented 39 of 40 directional laws, whereas lexical controls were near chance. Bidirectional interventions shifted answer probabilities toward or away from the physically appropriate outcome across all 12 matched cases, while counterfactual state patches transferred opposing decision signals across mechanisms and answer formats. Physical relationships were therefore more visible in controlled state changes than in absolute states alone.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!