2607.18100v1 Jul 20, 2026 cs.AI

LLM을 자기 루프에서 벗어나게 할 수 있을까? 활성화 제어를 통한 세밀한 추론 제어

Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering

Julian J. McAuley
Julian J. McAuley
Citations: 738
h-index: 15
Xunyi Jiang
Xunyi Jiang
Citations: 10
h-index: 2
Junda Wu
Junda Wu
Citations: 708
h-index: 15
Sheldon Yu
Sheldon Yu
Citations: 29
h-index: 3
Tong Yu
Tong Yu
Citations: 491
h-index: 11
Rohan Surana
Rohan Surana
Citations: 49
h-index: 5
Gagan Mundada
Gagan Mundada
Citations: 36
h-index: 4
Sungchul Kim
Sungchul Kim
Citations: 447
h-index: 10
Lina Yao
Lina Yao
Citations: 261
h-index: 10

확장된 추론은 최첨단 대규모 언어 모델(LLM)의 표준이 되었지만, 이러한 모델들이 생성하는 추론 경로는 여전히 대부분 통제하기 어렵습니다. 모델의 추론 방식을 형성하는 기존 방법들은 프롬프트 기반으로 작동하며 입력 수준에서만 제어가 가능하여, 자체적인 추론 과정에 대한 세밀한 제어를 제공하지 못합니다. 관련 연구에서는 대규모 언어 모델의 추론 경로에서 잠재적 전환 동역학을 분석하고 발견합니다. 본 연구는 이러한 상태를 통계적으로 특성화하고, 실패하는 추론 경로는 종종 자기 루프에 갇혀 최종 답변에 도달하지 못하면서 토큰 예산을 소진한다는 것을 보여줍니다. 이러한 실패를 해결하기 위해, 우리는 SOPHIA: Hidden-state Intervention 및 Activations 를 통한 추론 프로세스 제어(Steering Of reasoning Processes via Hidden-state Intervention and Activations) 방법을 제안합니다. 우리는 각 추론 경로를 비정형 텍스트가 아닌 잠재 상태의 시퀀스로 간주하고, 추론 시간 중 개입이 자기 루프에 빠지는 추론 과정을 세밀하게 제어할 수 있는지 조사합니다. 우리는 모든 접두사를 잠재 상태로 분류하고, 단계별 전환을 기록하며, 이를 사용하여 상태 쌍으로 인덱싱된 제어 벡터 집합을 구축합니다. 추론 시, 제어기는 현재 상태를 추론하고, 목표 상태가 주어지면 해당 벡터를 검색하며, 또한 전환 구조에서 자기 루프를 실시간으로 감지하여 모델이 추론의 블랙홀에 빠지는 것을 방지할 수 있습니다. 광범위한 실험을 통해, 우리의 방법은 자기 루프 실패에 대해 안정적으로 개입하며, 다양한 상태 쌍에 일반화되는 제어 벡터를 사용합니다. 최종 작업 정확도 및 토큰 효율성은 세밀한 제어가 더 나은 추론 품질로 이어진다는 것을 나타냅니다.

Original Abstract

Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models produce remain largely uncontrollable. Existing methods for shaping how a model reasons are prompt based approaches and operate at the input level, offering no fine-grained control over the reasoning process itself. Related work analyzes and discovers latent transition dynamics in the reasoning traces from Large Language Models. Building on this, we statistically characterize these states, and show that failure trajectories get stuck in self-loops, exhausting the token budget without progress toward the final answer. To intervene on these failures, We propose SOPHIA: Steering Of reasoning Processes via Hidden-state Intervention and Activations. We treat each reasoning trace as a sequence of latent states rather than an unstructured texts, and investigate whether inference time interventions can provide fine-grained control over the self-looping reasoning process. We classify every prefix to a latent state, record step level transitions, and use them to construct a bank of steering vectors indexed by state pairs. At inference time, a controller infers the current state and, given a target state, retrieves the corresponding vector and can also detect self-loops online from the transition structure to prevent the model from sinking into a reasoning black hole. Through extensive experiments, our method reliably intervenes on self-loop failures, with steering vectors that generalize to different state pairs. End task accuracy and token efficiency indicate that fine-grained controllability results in better reasoning quality.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!