인간의 능력을 초월하는 지능 측정: 인간의 한계를 넘어
Measuring Intelligence Beyond Human Scale
어떻게 인간의 능력 범위를 뛰어넘는 지능을 측정할 수 있을까요? 현재 사용되는 평가 기준은 인간 수준에 도달한 것으로 보이며, 인간의 능력을 넘어서는 경우, 평가자조차 어떤 작업이 어렵고 검증 가능한지 판단하기 어려울 수 있습니다. 우리는 이러한 어려움이 절대적인 기준으로 인한 것이라고 주장하며, 모델들이 다른 시스템과 차별성을 보여주는 공개적인 과제를 생성하는 상대적 측정 방식을 기반으로 하는 새로운 패러다임을 제안합니다. 이러한 결과들을 종합하여, 측정 대상 시스템의 규모에 맞춰 확장 가능한 적대적 심리측정 평가 시스템을 구축할 수 있습니다. 우리는 개인 정보 공격에 대한 유인을 줄이고, 판사 없이도 공정한 판단이 가능하도록 하며, 에이전트의 능력과 함께 자연스럽게 확장되는 실용적인 프로토콜을 제시합니다. 제안하는 프레임워크를 검증 가능한 영역과 비검증 가능한 개방형 영역에서 구현하여, 모델이 생성한 평가가 인간의 한계를 넘어 시스템을 측정하는 방법을 보여줍니다.
How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not know which tasks are both hard and verifiable. We argue that this difficulty is inherent to absolute-scale evaluation and propose a new paradigm based on relative measurement in which models generate public challenges that separate other systems. Aggregating these outcomes yields an adversarial psychometric rating system that can scale with the systems being measured. We describe practical protocols that reduce incentives for private-information attacks, support judge-free adjudication, and naturally scale with agent capabilities. We instantiate the framework across verifiable and open-ended, non-verifiable domains, illustrating how model-generated evaluation can continue to measure systems beyond the human frontier.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.