2607.26928v1 Jul 29, 2026 cs.CL

Latent-IM: 음성 LLM을 위한 잠재 상호작용 관리

Latent-IM: Latent Interaction Management for Speech LLMs

Larry Heck
Larry Heck
Citations: 46
h-index: 4
Atahan Dokme
Atahan Dokme
Citations: 1
h-index: 1
Adar Avsian
Adar Avsian
Georgia Institute of Technology
Citations: 17
h-index: 2
Tony Woo
Tony Woo
Citations: 0
h-index: 0

전통적인 음성 대화 시스템은 일반적으로 대화 관리를 응답 생성과 분리하여 처리했습니다. 즉, 정책(policy)이 다음 대화 동작을 선택하고, 생성 구성 요소가 해당 동작을 표현합니다. 그러나 대화 시스템이 LLM으로 전환되면서 이러한 분해는 모델의 숨겨진 표현 안에 대부분 사라졌습니다. 본 연구에서는 LLM 내부에서 상태 추정과 행동 제어에 해당하는 기능을 회복하여, 긍정 응답(acknowledging), 확인(checking), 질문(querying), 설명(explaining), 답변(replying)과 같은 대화적 움직임을 수행할 수 있는지 탐구합니다. 우리는 이러한 움직임 제어를 선택(selection) 및 실현(realization)이라는 두 가지 연결된 문제로 정의했습니다. 여기서 선택은 대화 맥락에서 적절한 다음 움직임을 예측하는 것이고, 실현은 생성 시에 선택된 움직임을 자연스럽게 생성하는 것입니다. 본 연구에서는 Latent-IM이라는 내부 대화 관리 프레임워크를 제시합니다. 이 프레임워크는 다양한 목표 하에서 대화적 움직임을 선택하고 적용하기 위한 일반적인 인터페이스를 제공합니다. 본 연구에서는 이러한 제어를 사용하여 인간의 대화 움직임을 재현하며, 제어되지 않은 기본 모델(unsteered backbone)에 비해 평균 엔드투엔드 움직임 정확도를 12.5 포인트 향상시켰습니다. 성능은 미세 조정(fine-tuning)과 유사한 수준입니다.

Original Abstract

Classical spoken dialogue systems often separated dialogue management from response realization: a policy selected the next dialogue action, and a generation component expressed that action. As dialogue systems shift toward LLMs, this decomposition has largely disappeared into the model's hidden representations. We ask whether an LLM-internal analogue of state estimation and action control can be recovered for conversational moves such as acknowledging, checking, querying, explaining, and replying. We formulate move control as two coupled problems: selection, predicting the appropriate next move from the dialogue context, and realization, causally producing a chosen move at generation time. We introduce Latent-IM, an internal dialogue-management framework that provides a general interface for choosing and deploying conversational moves under different objectives. Here, we use this control to reproduce human move choices, improving average end-to-end move accuracy by 12.5 points over the unsteered backbone while performing comparably to fine-tuning.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!