SocietyBench: 반사실적 사회 현상 변화 예측
SocietyBench: Forecasting Counterfactual Social-World Evolution
최근 대규모 언어 모델(LLM)과 이를 기반으로 구축된 에이전트들은 특정 작업을 수행할 수 있는지 여부, 즉 버그 수정, 브라우저 제어, GUI 작동 등 능력에 대한 벤치마크 테스트를 많이 받고 있습니다. 하지만 실제 사회 현상이 어떻게 전개되는지를 얼마나 잘 이해하고 예측하는지, 즉 모델의 사회적 역량은 거의 측정되지 않았습니다. 본 논문에서는 SocietyBench라는 통합 벤치마크를 소개합니다. SocietyBench는 한 줄로 표현된 사건 주제를 입력받아, 다섯 개의 플랫폼에서 수집한 웹 뉴스 및 소셜 미디어 게시물을 활용하여 날짜별 순서대로 정리하고, 사실 정보와 여론을 분리한 타임라인을 구축합니다. 그런 다음, 타임라인 상의 각 특정 시점을 예측 질문 은행으로 구성합니다. 질문은 확률 정확성(probability calibration)과 시간적 정확성(temporal accuracy)이라는 서로 다른 두 가지 100점 척도로 평가됩니다. 모델이 타임라인을 보기 전에, 세 단계에 걸친 과정을 통해 모든 명칭 개체를 변경하고 각 사건별로 날짜를 일정 값만큼 이동시켜 실제 사건을 반사실적인 사회 환경으로 변환합니다. 즉, 구조는 동일하지만 모델이 사전 학습된 메모리와 비교하여 일치시킬 수 있는 표면 레이블은 제거됩니다. 중국어 및 영어 버전의 5가지 다양한 사건과 125개의 예측 지점에 대한 실험 결과, 가장 뛰어난 6개의 최첨단 LLM 중 최고 성능을 보이는 모델도 100점 만점에 75.0점을 기록했으며, 이는 단순한 기준선인 50점을 훨씬 웃도는 수치입니다. 두 가지 평가 척도는 독립적으로 작용합니다. 즉, 모델은 확률 정확성에서는 강점을 보이지만 시간적 정확성에서는 약점을 보이거나, 그 반대의 경우도 나타납니다. 공유 기반 모델을 사용하는 세 가지 에이전트 프레임워크는 기본 모델보다 성능이 향상되지 않았으며, 모델에 의존하지 않는 두 가지 휴리스틱 방법은 모든 LLM보다 성능이 떨어졌습니다. 사건별로 평가 지점 간의 격차는 최대 21.4점에 이르렀으며, 이는 여러 사건을 대상으로 평가하는 것이 하나의 사건만을 대상으로 평가하는 것보다 더 의미 있다는 우리의 주요 주장을 뒷받침합니다. 모든 익명화된 타임라인, 질문 은행, 정답 데이터 및 평가 코드는 공개됩니다.
Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how well a model understands and forecasts the way real social events unfold, has barely been measured. We introduce SocietyBench, an end-to-end benchmark that takes a one-line event topic, collects Web news and social-media posts across five platforms, distills them into a date-indexed timeline that keeps factual events and a public-opinion layer separate, and then turns every cutoff date on that timeline into an audited bank of forecasting questions. Questions are scored on two orthogonal 100-point axes: probability calibration and temporal accuracy. Before any model sees a timeline, a three-phase procedure replaces every named entity and shifts every date by a per-event constant, turning a real arc into a counterfactual social world -- structurally identical to what happened, but stripped of the surface labels a model could match against pre-training memory. On five heterogeneous events and 125 prediction points in Chinese and English editions, the strongest of six frontier LLMs reaches only 75.0 out of 100, against a trivial anchor of 50. The two axes come apart: a model can be calibration-strong but time-weak, or the reverse. Three agent frameworks built on a shared base model fail to improve on that base, and two model-free heuristics trail every LLM. Per-event gaps reach 21.4 points on a single axis, which is our main argument for evaluating on several events rather than one. All anonymized timelines, question banks, ground truth, and scoring code are released.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.