2606.12289v1 Jun 10, 2026 cs.LG

표준 해석 가능 모델: 라그랑주 역학을 기반으로 해석 가능한 방법을 연역적으로 설계하는 일반 이론

The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics

M. Jamnik
M. Jamnik
Citations: 3,049
h-index: 25
Pietro Barbiero
Pietro Barbiero
Citations: 654
h-index: 12
M. Zarlenga
M. Zarlenga
Citations: 518
h-index: 11
Francesco Giannini
Francesco Giannini
Citations: 30
h-index: 3
Giuseppe Marra
Giuseppe Marra
Citations: 68
h-index: 3
F. Bonchi
F. Bonchi
Citations: 39
h-index: 3
G. Felice
G. Felice
Citations: 690
h-index: 12
R. Noris
R. Noris
Citations: 90
h-index: 5

인공지능 모델의 복잡성이 증가함에 따라, 해석성은 모델의 작동 방식을 이해하고, 오류를 수정하며, 제어하기 위한 필수적인 도구가 되었습니다. 그러나 해석 가능성을 통해 해석 가능한 방법을 연역적으로 설계할 수 있는 일반적인 이론은 부족합니다. 이러한 이론과 방법 간의 격차는 분절된 연구 결과와 일관성 없는 평가 프로토콜로 이어집니다. 이러한 격차를 해소하기 위해, 라그랑주 역학에 기반한 일반 이론인 표준 해석 가능 모델(Standard Interpretable Model, SIM)을 제안합니다. SIM은 목표 사용자에 대한 해석 가능성의 정의를 일련의 전제로 요약하며, 이러한 전제로부터 해석 가능성 대칭과 이에 상응하는 제약을 체계적으로 도출합니다. 이들은 라그랑주 함수 경관을 형성하며, 이 함수의 최소값은 최적의 해석 가능한 모델에 해당합니다. 최소값을 찾기 위해, 기존의 불투명한 모델의 파라미터 값을 조정하여 더 해석 가능하게 만들거나, 제약을 해석 가능한 아키텍처로 컴파일할 수 있습니다. 실험적으로 SIM이 기존 방법의 한계를 식별하고 해결하며(여기에는 전통적인 방법, 개념 기반 방법 및 메커니즘 해석 포함), 아직 탐구되지 않은 연구 방향을 강조하고, 핵심 프로그래밍 인터페이스 설계를 위한 정보를 제공한다는 것을 보여줍니다. SIM은 단순한 연구 방법일 뿐만 아니라, 연역적 특성을 통해 해석 가능성 교육 과정의 토대를 제공하며, 오랫동안 분절되어 온 학문 분야에 대한 과학 공동체의 관점을 변화시킬 수 있습니다.

Original Abstract

As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their computations. However, interpretability lacks general theories to deductively design interpretable methods. This gap between theories and methods results in a fragmented literature and inconsistent evaluation protocols. To fill this gap, we introduce the Standard Interpretable Model (SIM), a general theory grounded in Lagrangian mechanics that enables the deductive design of interpretable methods. Specifically, the SIM summarises, in a set of premises, what interpretability is for a target user. From these premises, the SIM systematically derives interpretability symmetries and corresponding constraints, which shape the landscape of a Lagrangian whose minima correspond to optimal interpretable models. To reach the minima, one can either update the parameter values of an opaque model to make it more interpretable or compile constraints into an interpretable architecture. We empirically show that the SIM identifies and solves limitations of existing methods (including traditional, concept-based, and mechanistic interpretability), highlights underexplored research directions, and informs the design of core programming interfaces. Beyond being a research method, the deductive nature of the SIM offers pedagogical grounding for interpretability curricula and may shift the scientific community's perspective of a discipline that has long been fragmented.

0 Citations
0 Influential
12.5 Altmetric
62.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!