2604.21827v1 Apr 23, 2026 cs.AI

정렬(Alignment)에는 '판타지아 문제'가 있다

Alignment has a Fantasia Problem

Nathanael Jo
Nathanael Jo
Citations: 8
h-index: 2
Zoe De Simone
Zoe De Simone
Citations: 7
h-index: 2
Mitchell Gordon
Mitchell Gordon
Citations: 307
h-index: 4
Ashia Wilson
Ashia Wilson
Citations: 94
h-index: 6

현대의 AI 어시스턴트는 사용자가 목표를 명확하게 표현하고 필요한 도움을 구하도록 가정하며, 사용 지침을 따르도록 훈련됩니다. 그러나 수십 년에 걸친 행동 연구에 따르면, 사용자는 종종 목표가 완전히 형성되기 전에 AI 시스템과 상호 작용하는 경우가 많습니다. AI 시스템이 프롬프트를 의도의 완전한 표현으로 간주할 때, 이는 유용하거나 편리해 보일 수 있지만, 반드시 사용자의 필요에 부합하는 것은 아닙니다. 우리는 이러한 실패를 '판타지아 상호 작용'이라고 부릅니다. 우리는 판타지아 상호 작용이 정렬 연구에 대한 재고를 요구한다고 주장합니다. 즉, AI는 사용자를 합리적인 존재로 취급하는 대신, 사용자가 시간에 따라 의도를 형성하고 개선하도록 적극적으로 지원함으로써 인지적 지원을 제공해야 합니다. 이는 머신 러닝, 인터페이스 디자인, 행동 과학을 융합하는 학제간 접근 방식을 필요로 합니다. 우리는 이러한 분야의 통찰력을 종합하여 판타지아 상호 작용의 메커니즘과 실패를 분석합니다. 그런 다음 기존의 개입 방법이 왜 충분하지 않은지 보여주고, 인간이 작업 과정에서 느끼는 불확실성을 더 잘 해결하는 AI 시스템을 설계하고 평가하기 위한 연구 과제를 제안합니다.

Original Abstract

In accomplishing complex tasks, human cognition typically progresses from abstract to concrete (e.g., from brainstorming ideas to writing an essay). With the advent of highly capable AI assistants, people now offload various parts of their task to the system. However, instruction-tuned AI systems, even when designed to infer implicit intent, lack an understanding of humans' cognitive processes. When a user approaches AI while their goals and intentions are abstract, AI systems often short circuit their cognitive process through which those goals would be refined by jumping toward a final output (e.g., writing the essay entirely). Doing so takes away the user's agency in achieving the task: they may need to spend more time revising or, worse, settle on a suboptimal outcome. We call these failures Fantasia interactions after the famous Disney scene. We argue that Fantasia interactions demand a rethinking of alignment research, where AI systems optimize how cognitive responsibility is allocated within an interaction. We highlight gaps in state-of-the-art alignment methods, and outline a research agenda for training and evaluating models to achieve this vision.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!