2607.25132v2 Jul 27, 2026 cs.LG

희소 오토인코더를 이용한 해석 가능한 GOHR 에이전트

Interpretable GOHR Agents via Sparse Autoencoders

Hao Wang
Hao Wang
Citations: 31
h-index: 3
Jacob Feldman
Jacob Feldman
Citations: 6
h-index: 2
Yusong Zhao
Yusong Zhao
Citations: 0
h-index: 0
Shiwei Tan
Shiwei Tan
Citations: 73
h-index: 4
Weiyi Qin
Weiyi Qin
Citations: 347
h-index: 3
Wentian Wang
Wentian Wang
Citations: 4
h-index: 1
L. Gallos
L. Gallos
Citations: 5,105
h-index: 23
Paul B. Kantor
Paul B. Kantor
Citations: 4
h-index: 1
Vladimir Menkov
Vladimir Menkov
Citations: 231
h-index: 7

학습된 의사 결정 시스템의 해석 가능성을 확보하는 핵심적인 과제는 해당 시스템 내부 표현에 행동을 설명하는 데 도움이 되는 개념들이 포함되어 있는지 확인하는 것입니다. 본 연구에서는 Game of Hidden Rules (GOHR)에서 토큰화된 자기 회귀 변환기 에이전트에 대한 해석 가능성 실험 결과를 보고합니다. 우리는 두 개의 숨겨진 규칙으로 구성된 간단한 과제를 중심으로 분석하는데, 이 과제에서 두 가지 숨겨진 규칙 모두 객체의 모양을 목표 버킷에 매핑하지만, 서로 다른 순열(permutation)을 사용합니다. 에이전트는 이러한 두 가지 숨겨진 규칙에서 추출된 에피소드를 사용하여 학습되며, 고정된 가중치로 평가됩니다. 에이전트에게는 명시적인 규칙 레이블이 제공되지 않으며, 명시적인 규칙 분류기를 사용하지 않습니다. 따라서 모든 규칙 정보는 상호 작용 기록으로부터 암묵적으로 추론해야 합니다. 이러한 설정에서, 올바른 규칙은 에이전트가 유용한 움직임을 시도하고 수락/거부 피드백을 관찰하기 전에는 식별할 수 없습니다. 에이전트의 의사 결정 토큰 임베딩에 대해 학습된 희소 오토인코더(SAE)는 이러한 구조를 복원합니다. 보류된 의사 결정들을 간단한 개념 (예: 선택된 모양 또는 버킷)으로 레이블링했을 때, 특정 개념에 대해 높은 선택성을 갖는 SAE 차원은 해당 개념이 존재하는 대부분의 의사 결정에서 활성화됩니다. 또한 개별적인 SAE 차원은 탐색적 규칙 가설을 테스트하고 부정적인 피드백 후에 전환하는 것과 같은 해석 가능한 전략과 상응합니다.

Original Abstract

A central challenge in interpreting learned decision-making systems is to determine whether their internal representations contain concepts that help explain their behavior. We report interpretability experiments for a tokenized autoregressive Transformer agent in the Game of Hidden Rules (GOHR). We focus on a compact two-rule task in which both hidden rules map object shapes to target buckets, but with different permutations. The policy is trained on episodes sampled from these two hidden rules and then evaluated with fixed weights. It is never given a rule label and does not use an explicit rule classifier; any rule information must be inferred implicitly from interaction history. In this setting, the correct rule is not identifiable before the agent tries an informative move and observes accept/reject feedback. Sparse autoencoders (SAEs) trained on the agent's decision-token embeddings recover this structure. When held-out decisions are labeled by simple concepts such as the chosen shape or bucket, SAE dimensions that are highly selective for a concept cover most decisions where that concept is present. Individual SAE dimensions also correspond to interpretable strategies such as probing one rule hypothesis and switching after negative feedback.

0 Citations
0 Influential
11.5 Altmetric
57.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!