장기 측정: 인간-AI 상호작용에 대한 종단적 이해를 향하여
Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
언어 모델은 '인간성'과 사용자의 일상생활로의 빠른 통합이라는 특징으로 인해 매우 새로운 유형의 기술로서 자리 잡았습니다. 이러한 특징들의 조합은 인지, 발달 및 사회-정서적 변화와 같은 장기적인 위험을 초래할 수 있으며, 이는 단기간 상호작용에서는 드러나지 않지만 사용자에게 지속적인 영향을 미칠 수 있습니다. 이것이 자연어 처리(NLP) 분야의 중요한 새로운 과제의 기반이 됩니다. 즉, 텍스트 생성에 대한 정적이고 단기적인 평가에서 벗어나 행동 변화를 장기적으로 측정하고 인간과 모델 간의 상호작용을 시간에 따른 관점에서 이해하는 것입니다. 본 연구에서는 종단 데이터를 통해 발생하는 현상을 이해하는 데 중요한 사회과학 분야에서 사용되는 측정 방법을 활용합니다. 우리는 NLP 분야의 계산 방법론이 이러한 측정 방법과 결합될 필요가 있으며, 이는 인간-모델 상호작용의 장기적인 안전 위험을 이해하는 것뿐만 아니라 모델 개발 방향을 긍정적인 결과로 이끌도록 하는 데 도움이 되어야 한다고 논의합니다. 모델과의 상호작용에 따른 인간 행동 변화를 모델링할 수 있는 능력은 문제 행동을 사후적으로 감지하는 것이 아닌 온라인으로 미리 감지하는 것을 가능하게 하며, 이는 사용자에게 장기적인 위험을 완화하기 위한 프레임워크에서 활용되어야 합니다.
Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users' daily lives. This combination of features can introduce longitudinal risks---cognitive, developmental and socio-affective changes in humans---that might not surface in short-term interactions, but can have lasting long-term effects on users. This forms the basis of a critical new mission for NLP: to pivot from static, short-term evaluations of text generations to long-term measurements of behavioral changes, towards a diachronic understanding of human-model interactions. In this work, we draw from measurements used in social science fields that are crucial to understand emergent phenomena in longitudinal data. We discuss how computational methods in the field of NLP need to be combined with such measurements, not only to understand long-term safety risks of human-model interactions, but to help steer model development towards positive rather than negative outcomes for users. This ability to model human behavioral shifts as a function of model interactions can facilitate online rather than post-hoc detection of problematic behaviors, and should be leveraged in alignment frameworks to mitigate long-term risks in users.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.