2608.12036v1 Aug 12, 2026 cs.AI

메커니스트: 인공지능을 과학적 도구로 활용하여 지능의 작동 원리를 발견하는 연구

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Julian McAuley
Julian McAuley
Citations: 43
h-index: 3
Shuofei Qiao
Shuofei Qiao
Citations: 1,542
h-index: 15
Tat-Seng Chua
Tat-Seng Chua
Citations: 46
h-index: 3
Shumin Deng
Shumin Deng
Citations: 647
h-index: 7
Fei Shen
Fei Shen
Citations: 37
h-index: 3
Zhixiang Cui
Zhixiang Cui
Citations: 0
h-index: 0

인공지능 모델은 다양한 분야에서 놀라운 성공을 거두었지만, 그 능력의 근본적인 작동 원리와 잠재적인 위험성에 대한 이해는 여전히 부족합니다. 인공지능 개발이 더욱 빠르게 진행되고 자동화되면서, 작동 원리에 대한 탐색은 대부분 수동으로 이루어지고 있으며, 이는 모델의 기능과 인간의 이해 및 제어 능력 간의 격차를 심화시킵니다. 이러한 격차를 해소하기 위해, 우리는 인공지능을 과학적 도구로 활용하여 인공지능 지능의 근본적인 작동 원리를 자율적으로 발견하는 시스템인 '메커니스트'를 소개합니다. 자율적인 작동 원리 발견을 지원하기 위해, 약 13,000개의 논문을 포함하는 해석 가능성에 초점을 맞춘 지식 그래프를 구축하고 이를 26개 분야에 걸쳐 43백만 개의 논문으로 구성된 다학제 데이터베이스와 통합했습니다. 또한, 작동 원리 분석, 인과적 개입 및 검증을 위한 32개의 핵심 방법을 포함하는 라이브러리를 개발했습니다. Claude Code 및 기존의 AI 과학 시스템과 비교했을 때, 메커니스트는 더 가치 있는 작동 원리 가설을 생성하고 실험을 더욱 안정적으로 수행합니다. 또한, 메커니스트는 모델의 동작을 발견하는 단계에서부터 인공지능 모델을 설명하고 제어하는 단계로 발전하는 것을 보여줍니다. 구체적으로, 메커니스트는 과학 연구소에서의 예상치 못한 안전 위험을 밝혀내며, 안전해 보이는 학습 데이터를 통해 위험한 특성이 여러 영역으로 전이될 수 있음을 보여줍니다. 또한, 메커니스트는 믿음의 작동 원리 이론을 개발하여 모델이 어떻게 세계 지식을 표현하고, 신념을 형성하며, 타인의 신념을 추론하는지, 그리고 이러한 작동 원리가 사전 학습 과정에서 어떻게 발생하는지를 밝혀냅니다. 마지막으로, 메커니스트는 이러한 작동 원리에 대한 통찰력을 실제적인 개입 방법으로 전환하여 다양한 시나리오에서 모델 성능을 향상시키고 과학적 기반 모델이 특정 특성을 가진 DNA 서열을 생성하도록 유도합니다.

Original Abstract

AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. To support autonomous mechanistic discovery, we construct an interpretability-focused knowledge graph of approximately 13,000 papers and integrate it with a multidisciplinary database of 43 million papers spanning 26 fields. We further curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates more valuable mechanism hypotheses and executes experiments more reliably. Mechanist also demonstrates a progression from discovering model behaviors to explaining and controlling AI models. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer across modalities through apparently safe training data. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Finally, Mechanist translates these mechanistic insights into practical interventions that improve model performance across diverse scenarios and steer scientific foundation models toward generating DNA sequences with specified properties.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!