MatrAIx: 83억 개의 가상 에이전트를 활용한 세계 시뮬레이션
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
인공지능 시스템 및 디지털 제품에 대한 인간 평가에는 비용이 많이 들고, 속도가 느리며 확장성이 떨어집니다. 오프라인 평가는 더 높은 확장성을 제공하지만, 종종 인간의 다양성과 상호 작용 방식을 간과합니다. 따라서 우리는 다양한 사용자를 대상으로 인공지능 시스템 및 디지털 제품을 테스트하기 위한 대규모 시뮬레이션 사용자 평가 인프라인 MatrAIx를 소개합니다. MatrAIx는 세 가지 핵심 구성 요소로 이루어져 있습니다. 첫째, Persona 8B는 1,290개의 범주형 차원으로 표현된 83억 개의 가상 에이전트 기록을 포함합니다. 이러한 기록은 상관 관계가 있는 속성을 유지하는 의존 그래프에서 샘플링하거나 인간이 작성한 프로필에서 파생됩니다. 우리는 약 1백만 개의 고품질 코어셋을 공개하며, 여기에는 599,847개의 실제 데이터 기반 가상 에이전트와 400,000개의 합성 데이터 기반 가상 에이전트가 포함되어 있습니다. 둘째, MatrAIx Playground는 다양한 사용자가 디지털 제품을 평가하고 상호 작용할 수 있는 네 가지 환경(설문 조사, AI 챗봇, 웹, 앱)을 제공합니다. 셋째, MatrAIx는 전자 상거래, 소프트웨어, 금융 및 의료를 포함한 25개 이상의 분야에 걸쳐 1,010개의 애플리케이션 작업을 지원합니다. 우리는 여덟 가지 대표적인 작업에 대해 총 18,189회의 평가 실험을 수행했습니다. 가상 에이전트는 Claude Opus 4.8, GPT 5.5 및 Claude Haiku 4.5의 세 가지 LLM으로 구동되었습니다. 결과적으로 얻어진 피드백은 가격 인상 후 망설임, AI 비서가 실패했을 때 계속 진행할 의향, 지연 허용 범위 등 가상 에이전트의 배경에 따른 의사 결정 및 선호도의 다양성을 보여줍니다. 우리는 두 가지 주요 검증 연구를 수행했습니다. 첫째, 400회의 통제된 실험을 통해 10가지 행동 속성에 걸쳐 네 가지 환경에서 가상 에이전트의 일관성을 평가했습니다. 선언된 행동은 366회(91.5%) 동안 표현되었거나 정확하게 억제되었습니다. 둘째, 인간 및 LLM 평가자는 실제 데이터 기반 가상 에이전트 추출 품질을 평가했습니다. 전반적으로 MatrAIx는 다양한 시뮬레이션된 사용자를 통해 인공지능 시스템 및 디지털 제품을 평가하기 위한 엔드 투 엔드 인프라를 제공합니다.
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.