2607.19949v1 Jul 22, 2026 cs.AI

SenWorld: 맥락 정보가 풍부한 평가 데이터를 생성하기 위한 디지털 트윈 시뮬레이션

SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data

Zenghui Zhou
Zenghui Zhou
Citations: 0
h-index: 0
Xiaoyang Li
Xiaoyang Li
Citations: 0
h-index: 0
Xiaoxuan Qiao
Xiaoxuan Qiao
Citations: 0
h-index: 0
Zhilang Wei
Zhilang Wei
Citations: 0
h-index: 0
Tianming Lei
Tianming Lei
Citations: 0
h-index: 0

스마트폰 개인 비서는 사용자의 장기적인 데이터를 기반으로 작동하지만, 이를 평가하려면 정답이 명확하게 알려진 맥락 정보가 풍부한 데이터가 필요합니다. 또한, 실제 기기 로그는 개인정보 보호 문제로 공유하기 어렵습니다. 이러한 문제를 해결하기 위해, 우리는 SenWorld라는 물리적으로 구체화되고 결정론적이며 이벤트 기반의 디지털 트윈 시뮬레이션을 제안합니다. SenWorld는 설계 과정에서 정답이 고정된 데이터를 생성하며, 사용자 프로필은 실제 지도, 날씨, 공휴일 및 네트워크 데이터를 기반으로 구축된 환경에서 하루를 보내게 됩니다. 모든 관측 가능한 신호는 전체 시스템 스냅샷에 기록되며, 각 평가 사례는 사후 주석이나 대규모 언어 모델(LLM) 심사관이 아닌 기존 레코드의 포인터를 통해 레이블링됩니다. 우리는 베이징에 거주하는 16명의 사용자 프로필을 사용하여 이 방법을 평가했습니다. 생성된 데이터는 실제 사용자의 벤치마크와 카테고리 분포(Jensen-Shannon divergence (JSD) 0.070), 그리고 일상적인 통신 기록의 패턴(JSD < 0.1)에서 유사한 결과를 보였지만, 생성된 레코드는 실제 레코드보다 짧습니다. 스크립트화되지 않은 상호 작용에도 불구하고, 사용자 프로필은 완전하게 상호 연결된 대화 서브그래프와 차별화된 행동 패턴을 형성합니다. 717개의 평가 사례를 분석한 결과, SenWorld는 생산 스마트폰 비서에서 78건의 오류를 발견했으며, 주로 통화 및 단문 메시지(SMS) 기록과 관련된 오류였으며, 연락처, 일정 및 알람 관련 오류는 발생하지 않았습니다. 스냅샷 포인터는 각 오류가 비서 측의 검색 오류임을 확인했으며, LLM 심사관은 전혀 사용되지 않았습니다. 전반적으로, SenWorld는 개인 정보 보호를 강화하고 재현 가능하며, 배포 과정을 검증할 수 있는 평가 데이터 생성 경로를 제공하며, 레이블은 설계 과정에서 고정됩니다.

Original Abstract

Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose correct answers are known, and real device traces are too privacy-sensitive to share. To address this challenge, we present SenWorld, a physically grounded, deterministic, event-sourced digital-twin simulation that generates such data with ground truth fixed by construction. In SenWorld, personas live through a full day in a world built from real map, weather, holiday, and network data; every observable signal is archived in full-system snapshots; and each evaluation case is labeled by a pointer to an existing record rather than by post-hoc annotation or a large language model (LLM) judge. We evaluate this method with 16 personas in Beijing. The generated data closely matches the held-out real-user benchmark in category distribution (Jensen--Shannon divergence (JSD) 0.070) and in the daily rhythm of communication records (JSD below 0.1), though generated records remain shorter than real ones. Without scripted interaction, personas form a fully reciprocated dialogue subgraph and differentiated behavioral repertoires. Projected into 717 evaluation cases, the generated data exposes 78 failures in a production smartphone assistant, concentrating on call and Short Message Service (SMS) records while contacts, schedules, and alarms never fail. The snapshot pointer confirms each failure as an assistant-side retrieval error, with no LLM judge involved. Overall, SenWorld offers a privacy-safe, reproducible, and distribution-checked path to evaluation data whose labels are fixed by construction.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!