CLAIM: 불확실성 측정을 통한 대규모 언어 모델의 능동적 명확화 기술
CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement
개방형 환경에서의 인간-컴퓨터 상호작용 시나리오에서, 대규모 언어 모델(LLM)은 종종 모호하거나 불완전한 사용자 쿼리를 접하게 됩니다. 이러한 경우, 직접적인 답변을 제공하는 것은 종종 과도하게 일반화되거나 오류가 있거나 정보량이 부족한 응답으로 이어질 수 있습니다. 반면, 명확화 질문을 하는 것은 상호작용의 품질을 크게 향상시킬 수 있습니다. 그러나 기존 접근 방식은 여전히 두 가지 근본적인 문제인 '명확화가 필요한 시점'과 '어떤 측면의 쿼리를 명확히 해야 하는지'를 해결하기 위해 수동으로 주석이 달린 데이터나 선호도 정렬에 크게 의존합니다. 이러한 의존성은 높은 주석 비용을 발생시키고 일반화 능력을 제한합니다. 이러한 문제점을 해결하기 위해, 우리는 개방형 환경에서 능동적 명확화 학습을 위한 불확실성 기반 프레임워크인 CLAIM을 제안합니다. CLAIM은 여러 모델 간의 답변 불일치로 인해 발생하는 엔트로피를 통해 쿼리 불확실성을 정량화함으로써 명시적인 인간 선호도 주석이 필요 없도록 합니다. 이 불확실성 신호는 고품질의 합성 데이터를 구축하는 데 사용되며, 이는 지도 학습과 강화 학습을 결합하여 통합된 명확화 결정 모델을 훈련할 수 있도록 합니다. 구체적으로, 우리는 엔트로피 기반의 불확실성 추정, 의미론적 클러스터링 및 추론 기반 판단을 통합하는 엔트로피 기반 합성 데이터 생성 파이프라인을 제안합니다. 이를 통해 명확화 요구 사항에 대한 안정적인 자동 주석이 가능합니다. CLAIM을 훈련하기 위해, 우리는 명확화 프로세스를 구조화된 의사 결정 생성 문제로 정의하고, 지도 미세 조정(SFT)과 그룹 상대 정책 최적화(GRPO)를 결합한 학습 패러다임을 채택했습니다. 실험 결과는 CLAIM이 수동으로 레이블링된 데이터에 의존하지 않고도 안정적이고 일반화 가능한 명확화 전략을 학습할 수 있음을 보여주며, 이는 LLM과의 실제 개방형 환경 상호작용에서 사전 이해를 위한 저렴하고 강력한 솔루션을 제공합니다.
In open-domain human-computer interaction scenarios, large language models (LLMs) frequently encounter user queries that are ambiguous or incomplete. In such cases, directly producing an answer often leads to overgeneralized, erroneous, or low-information responses. In contrast, asking clarifying questions can substantially improve interaction quality. However, existing approaches still rely heavily on manually annotated data or preference alignment to address two fundamental challenges: when clarification is necessary, and which aspect of the query should be clarified. This reliance incurs high annotation costs and limits generalization. To address these challenges, we propose CLAIM, an uncertainty-driven framework for active clarification learning in open-domain settings. CLAIM eliminates the need for explicit human preference annotations by quantifying query uncertainty through the entropy induced by answer disagreements across multiple models. This uncertainty signal is then used to construct high-quality synthetic data, enabling the training of a unified clarification decision model through a combination of supervised learning and reinforcement learning. Specifically, we propose an entropy-driven synthetic data generation pipeline that integrates entropy-based uncertainty estimation with semantic clustering and reasoning-based judgments, enabling reliable automatic annotation of clarification requirements. To train CLAIM, we formulate the clarification process as a structured decision generation problem and adopt a training paradigm that combines supervised fine-tuning (SFT) with group-relative policy optimization (GRPO). Experimental results demonstrate that CLAIM can learn stable and generalizable clarification strategies without relying on manually labeled data, offering a low-cost and robust solution for proactive understanding in real-world open-domain interactions with LLMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.