대화형 벤치마크
Interactive Benchmarks
표준 벤치마크는 포화 현상, 주관성, 그리고 일반화 성능의 부족으로 인해 점점 신뢰성이 떨어지고 있습니다. 우리는 모델의 지능을 평가하는 데 있어, 모델이 정보를 능동적으로 습득하는 능력을 평가하는 것이 중요하다는 점을 주장합니다. 우리는 대화형 벤치마크(Interactive Benchmarks)라는 통합적인 평가 패러다임을 제안합니다. 이 패러다임은 제한된 자원 하에서 모델의 추론 능력을 대화형 방식으로 평가합니다. 우리는 이 프레임워크를 두 가지 환경에서 구현했습니다. 첫째, 모델이 판사(judge)와 상호 작용하여 논리와 수학에서 객관적인 진실이나 답을 추론하는 '대화형 증명(Interactive Proofs)' 환경이고, 둘째, 모델이 장기적인 효용을 극대화하기 위해 전략적으로 추론하는 '대화형 게임(Interactive Games)' 환경입니다. 우리의 결과는 대화형 벤치마크가 모델의 지능을 견고하고 정확하게 평가하며, 대화형 시나리오에서 개선될 여지가 여전히 크다는 것을 보여줍니다. 프로젝트 페이지: https://github.com/interactivebench/interactivebench
Standard benchmarks have become increasingly unreliable due to saturation, subjectivity, and poor generalization. We argue that evaluating model's ability to acquire information actively is important to assess model's intelligence. We propose Interactive Benchmarks, a unified evaluation paradigm that assesses model's reasoning ability in an interactive process under budget constraints. We instantiate this framework across two settings: Interactive Proofs, where models interact with a judge to deduce objective truths or answers in logic and mathematics; and Interactive Games, where models reason strategically to maximize long-horizon utilities. Our results show that interactive benchmarks provide a robust and faithful assessment of model intelligence, revealing that there is still substantial room to improve in interactive scenarios. Project page: https://github.com/interactivebench/interactivebench
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.