SLALOM: 사회 시뮬레이션을 위한 시뮬레이션 라이프사이클 분석 - 장기 관찰 지표 활용
SLALOM: Simulation Lifecycle Analysis via Longitudinal Observation Metrics for Social Simulation
대규모 언어 모델(LLM) 에이전트는 생성적 사회과학 분야에 혁신적인 가능성을 제시하지만, 심각한 타당성 문제를 안고 있습니다. 현재 시뮬레이션 평가 방법론은 '멈춘 시계' 문제에 직면합니다. 즉, 시뮬레이션이 올바른 최종 결과를 도출했음을 확인하는 동시에, 그 결과에 도달하는 과정이 사회학적으로 타당했는지 여부는 간과합니다. LLM의 내부 작동 방식이 불투명하기 때문에, 사회 메커니즘의 '블랙 박스'를 검증하는 것은 지속적인 과제입니다. 본 논문에서는 SLALOM(Simulation Lifecycle Analysis via Longitudinal Observation Metrics)이라는 프레임워크를 소개합니다. SLALOM은 검증 방식을 결과 확인에서 과정 충실도로 전환합니다. 패턴 지향 모델링(POM)을 기반으로 SLALOM은 사회 현상을 특정 SLALOM 게이트, 즉 뚜렷한 단계를 나타내는 중간 지점 제약을 가진 다변량 시계열로 간주합니다. 동적 시간 워핑(DTW)을 사용하여 시뮬레이션 경로를 경험적 데이터와 정렬함으로써, SLALOM은 구조적 현실성을 평가하는 정량적 지표를 제공하여, 타당한 사회 역학을 무작위 잡음과 구별하고 보다 강력한 정책 시뮬레이션 표준을 구축하는 데 기여합니다.
Large Language Model (LLM) agents offer a potentially-transformative path forward for generative social science but face a critical crisis of validity. Current simulation evaluation methodologies suffer from the "stopped clock" problem: they confirm that a simulation reached the correct final outcome while ignoring whether the trajectory leading to it was sociologically plausible. Because the internal reasoning of LLMs is opaque, verifying the "black box" of social mechanisms remains a persistent challenge. In this paper, we introduce SLALOM (Simulation Lifecycle Analysis via Longitudinal Observation Metrics), a framework that shifts validation from outcome verification to process fidelity. Drawing on Pattern-Oriented Modeling (POM), SLALOM treats social phenomena as multivariate time series that must traverse specific SLALOM gates, or intermediate waypoint constraints representing distinct phases. By utilizing Dynamic Time Warping (DTW) to align simulated trajectories with empirical ground truth, SLALOM offers a quantitative metric to assess structural realism, helping to differentiate plausible social dynamics from stochastic noise and contributing to more robust policy simulation standards.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.