그래프 기반 증거 연결을 활용한 개인 맞춤형 심층 연구 질의 개선
Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence Grounding
사용자 요청은 심층 연구 에이전트의 연구 사양으로 작용하며, 어떤 증거를 찾고 어떻게 종합할지를 결정합니다. 개인 맞춤형 심층 연구에서 이러한 사양은 추가적으로 사용자의 목표, 제약 조건, 선호도 및 평가 기준을 반영해야 합니다. 사용자 컨텍스트는 심층 연구 파이프라인 내에 통합되거나, 입력으로 제공되는 연구 사양에 포함될 수 있습니다. 본 연구에서는 후자에 초점을 맞춰, 변경되지 않은 심층 연구 에이전트에 전달하기 전에 사용자 요청을 개인 맞춤형 연구 사양으로 개선합니다. 이는 세 가지 관련된 의사 결정을 필요로 합니다: 어떤 프레임 요소가 관련성이 있는지, 사용 가능한 사용자 컨텍스트가 이를 충분히 뒷받침하는지, 그리고 사용자 기억을 검색하거나 사용자에게 질문하거나 중단하고 질의를 다시 개선할 것인지 결정하는 것입니다. G-STEER는 학습 과정에서 프레임 요소를 의도 추출 그래프(Intent Elicitation Graph) 내의 추출 대상으로 구성하며, 이 그래프는 요소들 간의 의존성을 나타냅니다. G-STEER는 다양한 요소 의존성과 증거 조건을 포괄하는 그래프 기반 경로를 통해 명확화 정책을 학습합니다. 이 정책은 목표 범위와 증거 획득 비용 사이의 균형을 유지하면서 개선된 질의를 생성합니다. 실험 결과, G-STEER는 평가된 모든 심층 연구 에이전트에서 가장 높은 전반적인 가중치 목표 범위를 달성했으며, 강력한 기준 모델보다 사용자 질문 빈도가 약 1/3 수준으로 낮았습니다.
User requests serve as research specifications for deep research agents, shaping what evidence to seek and how to synthesize it. In personalized deep research, these specifications must additionally reflect user goals, constraints, preferences, and evaluation criteria. User context can be incorporated either within the deep research pipeline or into the research specification provided as its input. We focus on the latter, refining the user request into a personalized research specification before passing it to an unchanged deep research agent. This requires resolving three coupled decisions: which framing factors are relevant, whether the available user context sufficiently supports them, and whether to retrieve user memory, ask the user, or stop and refine the query. For training, G-STEER organizes framing factors as elicitation targets in an Intent Elicitation Graph that captures their dependencies. It learns a clarification policy from graph-scaffolded trajectories spanning diverse factor dependencies and evidence conditions. The policy produces a refined query while balancing target coverage against the costs of evidence acquisition. Experiments show that G-STEER achieves the strongest overall weighted target coverage and the highest downstream report personalization across both evaluated DRAs, while asking roughly one third as many user questions as a strong clarification baseline.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.