2608.06933v1 Aug 07, 2026 cs.CL

Ask-E: 모델의 성능에 따른 질문 생성 환경

Ask-E: An Environment for Calibrated Question Generation

Scott Geng
Scott Geng
Citations: 214
h-index: 3
Sarah M. Pratt
Sarah M. Pratt
Citations: 5
h-index: 1
Jae Sung Park
Jae Sung Park
Citations: 431
h-index: 9
Ali Farhadi
Ali Farhadi
Citations: 118
h-index: 5

현재, 우리는 모델을 훈련하고 평가하기 위해 모델의 능력 한계에 있는 문제들을 활용합니다. 이러한 문제를 만드는 것은 자체적으로 어려운 작업이며, 모델의 한계를 탐색하고 기존 질문 분포를 벗어난 일반화 능력을 요구합니다. 또한, 문제의 난이도를 정확하게 설정해야 하는데, 이는 문제를 해결하는 데 필요한 요소를 이해해야 합니다. 요약하자면, 모델의 현재 능력 한계에 맞춰 문제를 생성하려면 그 이상의 능력이 필요하며, 모델이 발전함에 따라 이러한 제약은 점점 더 부담스러워집니다. 우리의 핵심적인 통찰은 이 제약을 우리에게 유리하게 활용할 수 있다는 것입니다. 특정 능력 수준에 일관되게 맞춰진 문제를 생성하는 모델은 반드시 그 이상의 능력을 가지고 있어야 합니다. 따라서, 우리는 Ask-E를 소개합니다. Ask-E는 모델이 주어진 숙련도 수준에서 질문을 작성하는 능력을 평가하고 훈련하는 환경입니다. 구체적으로, 우리는 대상 숙련도 수준을 두 개의 기존 언어 모델의 능력으로 정의된 범위로 설정했습니다. 생성된 질문이 정확하게 하나의 모델에 의해 해결될 경우, 해당 질문은 대상 범위 내에 정확하게 배치되며, 이 두 모델의 능력을 구분할 수 있습니다. Ask-E는 다양한 숙련도 수준에 맞춰 문제를 생성하는 벤치마크이자 훈련 환경 역할을 합니다. 우리는 최첨단 모델조차 벤치마크에서 50% 미만의 정확도를 보이는 것을 확인했으며, 이는 향후 발전 가능성이 상당함을 보여줍니다. 또한, 이 환경에서의 훈련은 새로운 수학 데이터나 더 강력한 모델과의 상호 작용 없이, 그리고 정확도 기반 보상을 사용하지 않고서도 다양한 하위 수학 벤치마크에서 성능 향상을 가져온다는 것을 보여줍니다.

Original Abstract

Today, we improve models by training and evaluating them on problems at the frontier of their abilities. Creating such problems is itself a demanding task, requiring the ability to probe model limits and generalize beyond existing question distributions. It also means placing problems at a precise difficulty level, which requires understanding what it takes to solve them. In short, generating problems calibrated to a model's current frontier demands capability beyond it, an increasingly burdensome constraint as models improve. Our key insight is that we can leverage this constraint to our advantage: a model that can generate problems consistently calibrated to a given frontier must possess capability beyond it. Accordingly, we present Ask-E, an environment that benchmarks and trains models on their ability to write questions at a given skill level, rather than answer them. Concretely, we define target skill levels as ranges bounded by the capabilities of two existing language models. A generated question is successfully calibrated if exactly one of the two models can solve it, placing it precisely within the target range and differentiating the capabilities of these models. Ask-E serves both as a benchmark and a training environment, where models generate problems calibrated to a variety of skill levels. We find that even frontier models achieve below 50% calibration on the benchmark, leaving significant headroom to measure future progress. We also show that training on this environment leads to improvements across a number of downstream math benchmarks even with no new math data, no interaction with stronger models, and no correctness-based reward.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!